Artificial Intelligence Engineer
Role Overview
We are looking for an AI Engineer with strong AI/ML, GenAI and agentic systems expertise to build and productionize next-generation AI systems. This role is for someone who goes beyond using frameworks: someone who understands how models, retrieval systems and agents work internally, can design the architecture and orchestration around them, evaluate them rigorously, and take a prototype all the way to a scalable, reliable, production-grade system.
The candidate will lead the following workstreams:
1. AI/ML & Multimodal AI
- Strong first-principles understanding of Machine Learning, Deep Learning and Computer Vision
- Deep understanding of LLMs and Vision-Language Models (VLMs): tokenization/encoding, vision and language representations, multimodal alignment and how semantic information is fused
- SFT for LLMs/VLMs, PEFT/LoRA and fine-tuning strategies, including fine-tuning for tool use, function calling and structured outputs
- Model optimization: quantization (AWQ, INT4, HQQ), distillation, batching and inference optimization
- vLLM and high-performance model serving, including guided/constrained decoding and serving models behind agent workloads
- Reward modelling, custom/verifiable rewards, DPO, RLHF and GRPO, including reinforcement learning on multi-step, tool-using agent trajectories
- Hands-on experience with PyTorch, Hugging Face Transformers and TRL
2. RAG, Retrieval & Knowledge Systems
- Build advanced RAG / GraphRAG / VisionRAG systems, including agentic RAG where an agent plans, routes and iterates over retrieval steps
- Retrieval and ranking: BM25, semantic retrieval, hybrid search, re-ranking (cross-encoders, LLM re-rankers) and vector databases
- Understand and implement RRF, MRR, Recall, Precision, F1, RAGAS and other retrieval evaluation approaches
- Query optimization using HyDE, query expansion, decomposition and rewriting
- Document ingestion and chunking strategies for text, tables, images and scanned/handwritten documents
- Knowledge graph construction and entity/relationship extraction to power GraphRAG
- Expose retrieval and knowledge stores as reusable tools and MCP servers that agents can call
3. Agentic Systems & Orchestration
Design, build and run reliable single- and multi-agent systems that plan, use tools and complete multi-step business workflows in production. Agent architecture and design patterns
- Judge when a problem needs an autonomous agent versus a deterministic workflow or a simple LLM call, and design accordingly
- Single- and multi-agent patterns: ReAct, plan-and-execute, router, orchestrator–worker,
supervisor/hierarchical, handoff/swarm and evaluator–optimizer (reflection, self-critique)
- Task decomposition, planning and reasoning strategies, including effective use of reasoning models and extended thinking
- Long running and asynchronous agents, background tasks and agents that operate over hours or days Orchestration frameworks and runtimes
- Hands-on with LangGraph (state graphs, checkpointing, interrupts, sub-graphs) and at least one of CrewAI, AutoGen/AG2, OpenAI Agents SDK, Claude Agent SDK, Google ADK, Agno or equivalent
- Durable execution and workflow engines (Temporal or equivalent) for retries, resumability and exactly once side effects
- Coding, browser and computer-use agents, with sandboxed code execution
Tools, protocols and context engineering
- Build and consume MCP servers and clients: tools, resources and prompts; stdio and Streamable HTTP transports; authentication and authorization
- Agent-to-agent communication using A2A or equivalent protocols
- Tool design: clear tool descriptions, schemas, idempotent tools, timeouts, retries, fallbacks and error recovery
- Structured output generation, schema validation (Pydantic/JSON Schema) and reliable tool execution
- Context engineering: what goes into the context window and when, dynamic tool selection, prompt caching, summarization and compaction for long sessions State and memory
- Short-term/working memory and agent state management across steps and sessions
- Long-term memory (episodic, semantic, procedural), memory stores and retrieval of past interactions
- State persistence, checkpointing and resume-after-failure Human-in-the-loop, guardrails and safety
- Approval gates, interrupt-and-resume, escalation to humans and confidence-based routing
- Input/output guardrails, prompt-injection defense, PII handling and policy enforcement
- Least-privilege tool permissions, sandboxing, and controls on cost, step count and runaway loops
Agent evaluation and observability
- Evaluate agents end to end: task success rate, trajectory quality, tool-call accuracy, step efficiency, cost and latency per task
- LLM-as-judge, simulated users and environments, and regression suites for agent behaviour
- Tracing and observability with Open Telemetry, LangSmith, Langfuse, Arize Phoenix or equivalent; root-cause analysis of agent failures Scaling agents in production
- Model routing and fallbacks across providers, rate-limit handling and cost optimization
- Concurrency, queue-based agent execution and multi-tenant isolation
4. Evaluation & ML Engineering
- Build evaluation harnesses, automated testing and benchmarking frameworks for models, RAG pipelines and agents
- Design golden datasets, taxonomies and evaluation datasets, including multi-turn and multi-step agent task suites
- Perform model and agent benchmarking, error analysis and root-cause analysis of failures (wrong tool, bad plan, retrieval miss, hallucination)
- Run offline and online evaluation: A/B tests, canary releases and continuous evaluation on production traces
- Design data preprocessing, transformation and ML pipelines, including synthetic data generation for finetuning and evaluation
- Understand model quality vs. latency vs. memory vs. cost trade-offs, including token and step budgets for agents
- Experience with distributed ML systems using Ray or equivalent
5. Production AI Systems & Architecture
- Design end-to-end AI/ML system architectures, including agent runtimes, tool layers, memory stores and orchestration services
- Take rapid prototypes to robust production systems
- Build Python/FastAPI microservices and REST APIs, including streaming agent responses to front ends
- Docker/containerization, webhooks and SSE/WebSockets
- CI/CD, Git/GitHub and production engineering practices, with evaluation gates in the release pipeline for prompts, models and agents
- Design scalable pipelines involving queues, asynchronous processing, event-driven workflows and distributed workloads
- Deploy and operate MCP servers and agent services securely: secrets management, auth, rate limiting and audit logging
- Hands-on exposure to multi-GPU training/inference is highly desirable
Background
An ideal candidate will have the following background:
- Undergraduate degree in quantitative discipline such as engineering or science from a top-tier institution.
- Minimum 3 years of relevant experience in GenAI.
- Demonstrated experience designing, building and deploying agentic or multi-agent systems in production,with measurable business outcomes.
- Strong hands-on experience in AI/ML infrastructure, cloud platforms and production-grade ML systems.
- Prior experience with AWS or GCP cloud services, particularly AWS Bedrock (including Bedrock
Agents/AgentCore), SageMaker, EC2, S3, SQS, Lambda, Step Functions, ECR, EKS, CloudWatch, or GCP Vertex AI (including Vertex AI Agent Engine).
- Experience with GPU infrastructure, CUDA, multi-GPU environments and distributed training is required.
- Hands-on experience with containerized and Kubernetes-based environments, including Docker and Kubernetes.
- Experience working with LLMs, model training, fine-tuning, inference and deployment using modern ML frameworks and tools.
- Experience building and deploying scalable APIs, ML services and agent services in production environments.
- Strong understanding of Linux, networking/API fundamentals, REST APIs and CI/CD practices.
Skills
- Strong hands-on experience with Python, PyTorch, Hugging Face Transformers, TRL, FastAPI, Docker, vLLM and Ray.
- Hands-on experience with LangGraph and at least one other agent framework (CrewAI, AutoGen/AG2, OpenAI Agents SDK, Claude Agent SDK, Google ADK, Agno or equivalent).
- Practical experience building MCP servers/clients and integrating tools, APIs and enterprise systems into agents.
- Very good understanding of SQL/PostgreSQL, vector databases and Redis, including their use as agent state and memory stores.
- Hands-on experience with Kubernetes and distributed computing environments.
- Strong understanding of GPU/CUDA fundamentals, multi-GPU systems and distributed training.
- Experience with observability and monitoring tools, particularly Prometheus, Grafana, Loki, Promtail and CloudWatch, plus LLM/agent tracing tools such as LangSmith, Langfuse or Arize Phoenix.
- Experience with AWS/GCP cloud infrastructure and relevant AI/ML services.
- Strong understanding of Git/GitHub, REST APIs and CI/CD pipelines.
- Ability to design, build, deploy and monitor scalable AI/ML systems, agentic workflows and production inference services.
Location and rewards
Location: Chennai, hybrid work mode (4 days work from office, 1 day work from home)
Rewards:
- Attractive compensation
- Rapid and clear growth path
- Interaction with industry experts
- On-the-job training and skill development
Eucloid offers an expedited growth path along with a compensation package which is among the best in the industry.
Additional information: If you require reasonable accommodation during the application or selection process, please reach out to Eucloid HR at careers@eucloid.com.