Artificial Intelligence Engineer

Eucloid Data Solutions
Chennai, Tamil Nadu, India

Role Overview

We are looking for an AI Engineer with strong AI/ML, GenAI and agentic systems expertise to build and productionize next-generation AI systems. This role is for someone who goes beyond using frameworks: someone who understands how models, retrieval systems and agents work internally, can design the architecture and orchestration around them, evaluate them rigorously, and take a prototype all the way to a scalable, reliable, production-grade system.

The candidate will lead the following workstreams:

1. AI/ML & Multimodal AI

  • Strong first-principles understanding of Machine Learning, Deep Learning and Computer Vision
  • Deep understanding of LLMs and Vision-Language Models (VLMs): tokenization/encoding, vision and language representations, multimodal alignment and how semantic information is fused
  • SFT for LLMs/VLMs, PEFT/LoRA and fine-tuning strategies, including fine-tuning for tool use, function calling and structured outputs
  • Model optimization: quantization (AWQ, INT4, HQQ), distillation, batching and inference optimization
  • vLLM and high-performance model serving, including guided/constrained decoding and serving models behind agent workloads
  • Reward modelling, custom/verifiable rewards, DPO, RLHF and GRPO, including reinforcement learning on multi-step, tool-using agent trajectories
  • Hands-on experience with PyTorch, Hugging Face Transformers and TRL

2. RAG, Retrieval & Knowledge Systems

  • Build advanced RAG / GraphRAG / VisionRAG systems, including agentic RAG where an agent plans, routes and iterates over retrieval steps
  • Retrieval and ranking: BM25, semantic retrieval, hybrid search, re-ranking (cross-encoders, LLM re-rankers) and vector databases
  • Understand and implement RRF, MRR, Recall, Precision, F1, RAGAS and other retrieval evaluation approaches
  • Query optimization using HyDE, query expansion, decomposition and rewriting
  • Document ingestion and chunking strategies for text, tables, images and scanned/handwritten documents
  • Knowledge graph construction and entity/relationship extraction to power GraphRAG
  • Expose retrieval and knowledge stores as reusable tools and MCP servers that agents can call

3. Agentic Systems & Orchestration

Design, build and run reliable single- and multi-agent systems that plan, use tools and complete multi-step business workflows in production. Agent architecture and design patterns

  • Judge when a problem needs an autonomous agent versus a deterministic workflow or a simple LLM call, and design accordingly
  • Single- and multi-agent patterns: ReAct, plan-and-execute, router, orchestrator–worker,

supervisor/hierarchical, handoff/swarm and evaluator–optimizer (reflection, self-critique)

  • Task decomposition, planning and reasoning strategies, including effective use of reasoning models and extended thinking
  • Long running and asynchronous agents, background tasks and agents that operate over hours or days Orchestration frameworks and runtimes
  • Hands-on with LangGraph (state graphs, checkpointing, interrupts, sub-graphs) and at least one of CrewAI, AutoGen/AG2, OpenAI Agents SDK, Claude Agent SDK, Google ADK, Agno or equivalent
  • Durable execution and workflow engines (Temporal or equivalent) for retries, resumability and exactly once side effects
  • Coding, browser and computer-use agents, with sandboxed code execution

Tools, protocols and context engineering

  • Build and consume MCP servers and clients: tools, resources and prompts; stdio and Streamable HTTP transports; authentication and authorization
  • Agent-to-agent communication using A2A or equivalent protocols
  • Tool design: clear tool descriptions, schemas, idempotent tools, timeouts, retries, fallbacks and error recovery
  • Structured output generation, schema validation (Pydantic/JSON Schema) and reliable tool execution
  • Context engineering: what goes into the context window and when, dynamic tool selection, prompt caching, summarization and compaction for long sessions State and memory
  • Short-term/working memory and agent state management across steps and sessions
  • Long-term memory (episodic, semantic, procedural), memory stores and retrieval of past interactions
  • State persistence, checkpointing and resume-after-failure Human-in-the-loop, guardrails and safety
  • Approval gates, interrupt-and-resume, escalation to humans and confidence-based routing
  • Input/output guardrails, prompt-injection defense, PII handling and policy enforcement
  • Least-privilege tool permissions, sandboxing, and controls on cost, step count and runaway loops

Agent evaluation and observability

  • Evaluate agents end to end: task success rate, trajectory quality, tool-call accuracy, step efficiency, cost and latency per task
  • LLM-as-judge, simulated users and environments, and regression suites for agent behaviour
  • Tracing and observability with Open Telemetry, LangSmith, Langfuse, Arize Phoenix or equivalent; root-cause analysis of agent failures Scaling agents in production
  • Model routing and fallbacks across providers, rate-limit handling and cost optimization
  • Concurrency, queue-based agent execution and multi-tenant isolation

4. Evaluation & ML Engineering

  • Build evaluation harnesses, automated testing and benchmarking frameworks for models, RAG pipelines and agents
  • Design golden datasets, taxonomies and evaluation datasets, including multi-turn and multi-step agent task suites
  • Perform model and agent benchmarking, error analysis and root-cause analysis of failures (wrong tool, bad plan, retrieval miss, hallucination)
  • Run offline and online evaluation: A/B tests, canary releases and continuous evaluation on production traces
  • Design data preprocessing, transformation and ML pipelines, including synthetic data generation for finetuning and evaluation
  • Understand model quality vs. latency vs. memory vs. cost trade-offs, including token and step budgets for agents
  • Experience with distributed ML systems using Ray or equivalent

5. Production AI Systems & Architecture

  • Design end-to-end AI/ML system architectures, including agent runtimes, tool layers, memory stores and orchestration services
  • Take rapid prototypes to robust production systems
  • Build Python/FastAPI microservices and REST APIs, including streaming agent responses to front ends
  • Docker/containerization, webhooks and SSE/WebSockets
  • CI/CD, Git/GitHub and production engineering practices, with evaluation gates in the release pipeline for prompts, models and agents
  • Design scalable pipelines involving queues, asynchronous processing, event-driven workflows and distributed workloads
  • Deploy and operate MCP servers and agent services securely: secrets management, auth, rate limiting and audit logging
  • Hands-on exposure to multi-GPU training/inference is highly desirable

Background

An ideal candidate will have the following background:

  • Undergraduate degree in quantitative discipline such as engineering or science from a top-tier institution.
  • Minimum 3 years of relevant experience in GenAI.
  • Demonstrated experience designing, building and deploying agentic or multi-agent systems in production,with measurable business outcomes.
  • Strong hands-on experience in AI/ML infrastructure, cloud platforms and production-grade ML systems.
  • Prior experience with AWS or GCP cloud services, particularly AWS Bedrock (including Bedrock

Agents/AgentCore), SageMaker, EC2, S3, SQS, Lambda, Step Functions, ECR, EKS, CloudWatch, or GCP Vertex AI (including Vertex AI Agent Engine).

  • Experience with GPU infrastructure, CUDA, multi-GPU environments and distributed training is required.
  • Hands-on experience with containerized and Kubernetes-based environments, including Docker and Kubernetes.
  • Experience working with LLMs, model training, fine-tuning, inference and deployment using modern ML frameworks and tools.
  • Experience building and deploying scalable APIs, ML services and agent services in production environments.
  • Strong understanding of Linux, networking/API fundamentals, REST APIs and CI/CD practices.

Skills

  • Strong hands-on experience with Python, PyTorch, Hugging Face Transformers, TRL, FastAPI, Docker, vLLM and Ray.
  • Hands-on experience with LangGraph and at least one other agent framework (CrewAI, AutoGen/AG2, OpenAI Agents SDK, Claude Agent SDK, Google ADK, Agno or equivalent).
  • Practical experience building MCP servers/clients and integrating tools, APIs and enterprise systems into agents.
  • Very good understanding of SQL/PostgreSQL, vector databases and Redis, including their use as agent state and memory stores.
  • Hands-on experience with Kubernetes and distributed computing environments.
  • Strong understanding of GPU/CUDA fundamentals, multi-GPU systems and distributed training.
  • Experience with observability and monitoring tools, particularly Prometheus, Grafana, Loki, Promtail and CloudWatch, plus LLM/agent tracing tools such as LangSmith, Langfuse or Arize Phoenix.
  • Experience with AWS/GCP cloud infrastructure and relevant AI/ML services.
  • Strong understanding of Git/GitHub, REST APIs and CI/CD pipelines.
  • Ability to design, build, deploy and monitor scalable AI/ML systems, agentic workflows and production inference services.

Location and rewards

Location: Chennai, hybrid work mode (4 days work from office, 1 day work from home)

Rewards:

  • Attractive compensation
  • Rapid and clear growth path
  • Interaction with industry experts
  • On-the-job training and skill development

Eucloid offers an expedited growth path along with a compensation package which is among the best in the industry.

Additional information: If you require reasonable accommodation during the application or selection process, please reach out to Eucloid HR at careers@eucloid.com.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →