AI Tutor Architect (Founding Track)

Recollia
India

Most AI tutors are chatbots with a syllabus. Recollia is building the tutor medical students wish they'd had: one with a soul, running on our own model, cheap enough for every student to afford. You'll own it.

WHERE THIS GOES

This starts as an internship with a stipend of [Rs 20,000] and a possible incentive of [Rs 5,000] every two weeks, depending on performance. Make this the best AI tutor a medical student can use and it converts into a founding engineer role with equity. The bar is high on purpose: we want teaching craft, model training, and cost engineering in one person.

WHAT YOU'LL BUILD

  • The tutor's teaching behaviour and voice: how it finds the gap in a student's understanding, when it asks and when it tells, how it holds its character across a long session
  • Our own tutor model: open-weight, fine-tuned for medical education, served on our own infrastructure in place of most API calls
  • A cost architecture where most questions never reach a large model: caching, precomputed explanations, routing
  • A learner model on top of RAV, our unified-memory layer, plus the cases, quizzing, and spaced review around it
  • Evals that prove three things: students learned, the medicine is right, and the cheap path is as good as the expensive one

REQUIRED

Tutor craft

  • You've built an AI tutor that real students used, not a demo; you can show it and say where it failed
  • The difference between answering and teaching: Socratic questioning, scaffolding, graduated hints, misconception diagnosis
  • Learning science well enough to build on it: retrieval practice, spaced repetition, interleaving, mastery learning, and where each one stops working
  • A persona that holds up across a whole session: warm without flattery, honest when the student is wrong
  • Student modelling and memory: knowledge tracing, adaptive difficulty, what to remember about a student and what to let go

Our own model

  • Fine-tuning open-weight models for one profession: SFT, LoRA/QLoRA, preference tuning, and knowing which one a problem needs
  • Dataset work: training sets built from tutoring sessions and expert-reviewed medical content; spotting when the data, not the model, is what's wrong
  • Serving it ourselves: vLLM or llama.cpp, quantization and the quality it costs, TTFT, tokens/sec, cost per session against the API baseline
  • Knowing what a small model can't do yet, and routing those turns to a bigger one

Cost engineering

  • Caching at every layer: prompt caching, semantic caching, exact-match response caching; and knowing when a cached answer is the wrong answer for this student
  • Precomputation: thousands of students ask the same questions about the same syllabus, so most explanations should be generated once, reviewed, and served many times
  • No model call where none is needed: a review schedule is arithmetic, an answer check is a lookup
  • Routing in order of cost: cache first, small model next, large model last
  • You track cost per student the way others track latency

Evaluation

  • Teaching quality, not only answer correctness: rubrics, simulated students, LLM-as-judge and its failure modes
  • Proof that cheap holds up: our model against the API baseline, cached answers against fresh ones
  • Learning outcomes: retention a week later, performance on unseen questions, whether students come back
  • Regression suites for medical accuracy; you read transcripts every day

Medical education

  • No medical background required, but real curiosity about how medicine is learned: how a second-year gets through pharmacology, why clinical reasoning is hard to teach, what a good viva examiner does differently
  • You'll sit with students and doctors, watch them study, and learn enough medicine to know when the tutor is wrong

Engineering foundations

  • Python at production level: async, typing, tests, streaming APIs
  • FastAPI or equivalent, Postgres, Redis, Docker, Git, CI, GPU and CUDA troubleshooting
  • Prompt and context engineering across Claude, OpenAI, and open-weight models; tracing and experiment logs that show why version 12 beats version 11

No degree required. Skills over credentials.

WHAT HUNGER LOOKS LIKE TO US

  • You've used every AI tutor out there and can say exactly where each one falls short.
  • A bad transcript bothers you until it's fixed. So does a wasted API call.
  • You'd rather test with ten students than debate it.
  • You ask for more scope, not more instructions.
  • You say when something is broken, including our decisions.

HOW TO APPLY

Email careers@recollia.ai with the subject "AI Tutor" and one tutor or learning tool you built: repo, demo, or write-up. In three sentences, tell us what it does, what broke, and what you'd do differently. Then two more: what it cost to run, and what the best teacher you ever had did that no AI tutor does yet.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →