AI Platform Engineer
Location: Gurugram - office-based, Monday to Friday
Experience: 3-5 years total, with 2+ years in backend engineering
Reports To: Engineering Lead (individual contributor role)
Function: AI Platform Engineering | Agent Runtime | Evals
Position Overview
We are hiring an AI Platform Engineer to build and own the shared runtime that AI agents run on, and the evals that keep it honest. We run AI agents on behalf of brands and need one gateway to the models, one place where prompts get routed, and one gate that checks what is allowed before anything ships. They also need proof that they work, because a brand cannot take our word for it.
The Hard Problems
- Build one agent runtime that new use cases plug into, without forking the code
- Build evals strong enough to block a bad release, not just report on it
- Run a policy gate on every generated asset, fast enough to sit in the request path
- Keep many brands on one platform, where one brand's memory never reaches another brand's output
Core Responsibilities
- Agent runtime and LLM gateway: own one reusable platform for agent orchestration, tool execution and prompt routing, with a single path to Anthropic, OpenAI and other providers (caching, dedup, fallback, retries)
- Evals and release gates: build the offline and online eval harness, golden datasets, calibrated LLM-as-judge and a CI regression gate that stops a prompt change that drops a score
- Cost, retrieval and latency: own inference spend per unit of output, prompt caching, and vector search over brand history (pgvector or similar) wired into the agents
- Safety and brand isolation: own the sandboxed policy gate, disclosure metadata and audit log on every asset, prompt-injection defence, four autonomy tiers, a kill switch and one-step rollback
- Observability and design: trace every agent step with OpenTelemetry so a wrong output points to the step that caused it, and own high-level and low-level design
Required Qualifications
- 3-5 years of total experience, with 2+ years in backend engineering; Python and FastAPI in production
- Shipped an LLM or agent system to production (not a demo, not a notebook)
- Built a platform others build on: SDK, shared service or internal framework with real users
- Built an eval harness that blocked a bad release, and can describe the release it blocked
- Cost or latency work that can be quantified: cut a number and knows how
- RAG in production: embeddings, a vector store and defensible chunking choices
- Cloud and CI/CD: AWS or GCP, Docker, GitHub Actions or equivalent
- Distributed systems basics: caching, queues, idempotency, P99 latency
- Written design docs; led at least a small team or a large project end to end
Good to Have
- Eval tooling (Braintrust, LangSmith, DeepEval or your own)
- Policy engines (OPA/Rego, sandboxes)
- Content provenance (C2PA, watermarking, EU AI Act)
- Ad tech or martech experience
- Side projects with agents
Why Join
This is an early role in a company built from the ground up. You will build the platform every agent runs on and the measurement that proves it works, and what you build becomes the foundation others build on.