Senior FDE (AI & Backend)

Pragyaa
Gurugram, Haryana, India

Context

We are building AI-native products on the Anthropic API and MCP, backed by a Django/DRF + PostgreSQL platform with Celery/Redis pipelines and a React/Next.js frontend, deployed on Azure. The team works in an AI-assisted workflow (agentic AI coding tools are a daily tool), and the codebase is young: designs are still being set, and agent patterns, evals and observability are not yet standardised. This is an architect-and-build seat, not a keep-the-lights-on role. The Senior SDE is the technical authority on system design: they decide how services are cut, how data flows, how agents and connectors plug in, and how the system scales and fails safely. They set the bar the Junior SDEs work to.

What the role owns

Own the architecture and the services built on it: design the overall system (service boundaries, data model, integration and agent architecture, MCP tool surface), write the design docs and ADRs, then build the critical paths yourself, build reliable Django/DRF services and async pipelines, decide trade-offs on cost, latency and reliability, review code and designs, and mentor juniors so repeat issues drop.

The actual workload (by share of time)

System architecture & design (~25%)

Service boundaries, data flow, scalability, multi-tenancy, API contracts, design docs/ADRs, tech-debt strategy

  • e.g. "Design the platform for 10x document volume"
  • e.g. "Monolith vs services for the agent runtime"
  • e.g. "Write the ADR for the queue and storage choices"

Agent, skills & MCP connector architecture (~25%)

Custom agents/subagents, orchestration, skill design, connector and tool design, guardrails, context/memory

  • e.g. "Design the multi-agent flow for document review"
  • e.g. "Build the connector for the client's DMS"
  • e.g. "Turn this SOP into a reusable skill"

Backend delivery (~15%)

Django/DRF, PostgreSQL models and performance, JWT/OAuth2, permissions on the critical paths

  • e.g. "Redesign the data model for workflows"
  • e.g. "Index and query plan for the cases list"

Async pipelines & reliability (~15%)

Celery/Redis, retries, idempotency, LLM rate limits, evals, observability, SLOs

  • e.g. "LLM calls time out and double-process"
  • e.g. "Add regression evals and cost tracking"

Code review, mentoring & standards (~10%)

  • e.g. "Review a junior's agent PR"
  • e.g. "Write the AI-assisted-code review checklist"

DevOps, incidents & stakeholder scoping (~10%)

Docker/Azure/CI-CD, RCA, turning business asks into tickets

  • e.g. "Prod queue backed up — root-cause it"
  • e.g. "Break this client ask into scoped tickets"

Priority reality: stakeholders will call everything urgent and everything an AI problem. The Senior SDE re-triages against real business impact, pushes back when a problem needs no LLM, and escalates only with a diagnosis and a recommendation.

Tech stack

  • AI/LLM engineering: Anthropic API, MCP (Model Context Protocol), agentic AI coding tools (CLI/IDE), reusable agent skills (SKILL.md-style), custom agents/subagents, custom MCP connector creation, LLM agent architecture & orchestration, prompt engineering, multi-agent systems, AI workflow automation. Optional: OpenAI API, RAG pipelines.
  • Backend: Python, Django, Django REST Framework, PostgreSQL, Celery, Redis, JWT/OAuth2, Docker.
  • Architecture & system design: service/data architecture, API design, event-driven and async patterns, scalability, multi-tenancy, design docs/ADRs, observability.
  • DevOps/Tools: Azure, Git, CI/CD. Optional: Sentry, AWS.
  • Testing & Quality (optional): pytest, TDD.
  • Frontend: React/Next.js, TypeScript, REST API integration.

Preferred qualifications

Must-have

  • System design & architecture: has designed and evolved a non-trivial production system — service boundaries, data modelling, API design, caching, queues, consistency and failure modes — and can defend trade-offs in a design review.
  • Scalability & reliability: capacity planning, horizontal scaling, DB performance (indexing, query plans, partitioning), retries/idempotency, graceful degradation, SLOs.
  • Security & multi-tenant design: authN/authZ models, tenant isolation, secrets, audit logging, least-privilege access for agents and connectors.
  • Production backend: 4+ years in Python, with 2+ years on Django/DRF and PostgreSQL.
  • Shipped LLM features: at least one in production on the Anthropic API — tool use, prompt iteration, failure handling.
  • Custom MCP connector creation: designed, built and deployed MCP servers/connectors against real APIs — tool schemas, OAuth/auth, rate limits, error contracts, versioning.
  • Agent skills & custom agents: authored production skills (SKILL.md, references, scripts, evals) and configured custom agents/subagents with scoped tools and permissions.
  • Agent architecture: orchestration, multi-agent coordination, context/memory management, guardrails.
  • Async pipelines: Celery/Redis in production — retries, idempotency, queue design.
  • Auth & security: JWT/OAuth2, permissions, secrets handling.
  • Design communication: writes clear design docs/ADRs, runs design reviews and presents architecture to technical and non-technical stakeholders.
  • Delivery: Docker + CI/CD on Azure (or equivalent) with ownership of a live deployment.
  • Leadership: code review, design-doc writing and mentoring record; clear written comms with non-technical stakeholders.

Strongly preferred

  • Agentic AI coding tools in daily use, with a point of view on where AI-generated code needs extra review.
  • Event-driven & integration architecture: message queues/streams, webhooks, API gateways, versioned APIs, backward-compatible migrations.
  • Cloud architecture on Azure: networking, managed Postgres/Redis, containers/orchestration, cost-aware design, infrastructure as code.
  • Team enablement: has set conventions for skills, agent definitions and connector reviews so others can reuse them.
  • Evals & observability for LLM systems — regression sets, cost and latency tracking.
  • Business-workflow automation: turned messy real workflows into reliable agent flows.
  • React/Next.js + TypeScript enough to review and unblock frontend work.

Nice-to-have: OpenAI API and cross-model fallbacks; RAG pipelines (embeddings, pgvector, retrieval evals); pytest and TDD leadership; Sentry; AWS; PostgreSQL performance tuning; domain-driven design; distributed tracing (OpenTelemetry); Kubernetes; Terraform/Bicep.

This is first and foremost an architecture and AI-systems ownership hire. Weight evaluation toward system-design judgement (a design round), production LLM experience and the ability to raise the team's level — not a long list of tools.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →