AI Data Engineer (Backend Owner, Founding Track)

Recollia
India

Recollia builds AI tools for medical students. Behind every answer is a backend: the data that grounds it, the cache that makes it fast, the analytics that tell us whether it helped. We need one person to own all of it.

WHERE THIS GOES

This starts as an internship with a stipend of Rs 20,000 and a possible incentive of Rs 5,000 every two weeks, depending on performance. Build a backend that is fast, cheap to run, and never the reason we slow down, and it converts into a founding engineer role with equity. We're early, so the role grows as Recollia does: more scope, bigger systems, more say in how we build.

WHAT YOU'LL OWN

  • The backend architecture on AWS: services, APIs, data model, and the decisions behind them
  • The data pipelines our AI runs on: content ingestion, embeddings, training and eval sets, usage events
  • Latency: caching at every layer, so the product feels instant on a student's phone
  • Product analytics in PostHog: what we track, what it means, and what we change because of it
  • Reliability, security, and cost: uptime, student data, and the AWS bill are yours

REQUIRED

Backend and architecture

  • You've owned a production backend end to end; you can defend its design and name what you'd change
  • Python (FastAPI or equivalent) or TypeScript at production level: async, typing, tests
  • API design, Postgres schemas and indexes, queues and background jobs
  • Enough frontend (React or similar) to trace a slow screen back to the query behind it

AWS and cloud

  • Core AWS in depth: VPC and IAM, ECS/Fargate or Lambda, RDS or Aurora, S3, CloudFront, SQS
  • Infrastructure as code (Terraform or CDK), CI/CD, zero-downtime deploys, rollbacks
  • Observability: logs, metrics, tracing, and alerts that mean something
  • Cost: you read the AWS bill and know which line to cut

Latency and caching

  • Caching at every layer: CDN, Redis, query caches, LLM prompt and response caches
  • Invalidation done right; a fast stale answer is worse than a slow correct one
  • You measure p50 and p95, find the slow hop, and fix that one

Data for AI

  • Pipelines for ingestion, cleaning, dedup, and versioning; data you can trust and reproduce
  • Embeddings and vector stores (pgvector or similar) for retrieval
  • Turning product usage into training and eval data for our models
  • Privacy by default for student data: PII, consent, and retention under India's DPDP Act

Analytics

  • PostHog or equivalent: event design, funnels, retention, feature flags, experiments
  • You can answer "which feature do students come back for?" with data, and you've changed a plan because of an answer like that

No degree required. Skills over credentials.

WHAT HUNGER LOOKS LIKE TO US

  • You treat the backend as yours. When it breaks at 2 a.m., you want to know why before anyone asks.
  • You'd rather measure it than argue about it.
  • You ask for more scope, not more instructions.
  • You say when something is broken, including our decisions.

HOW TO APPLY

Email careers@recollia.ai with the subject "AI Data Engineer" and one system you built and ran: repo, architecture diagram, or write-up. In three sentences, tell us what it does, what broke, and what you'd do differently. Then one more: its p95 latency, and what you did to get it there.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →