AI Engineer III
Rekruton Global IT Services
Ahmedabad, Gujarat, India
What You'll Do
- Lead the design and implementation of complex AI systems spanning models, retrieval, tools, agents, data, APIs, and existing IQM services.
- Own AI capabilities end-to-end from technical discovery and architecture through deployment, production operations, and continuous improvement.
- Design scalable RAG and retrieval architectures, context and memory strategies, agent/tool workflows, and model-provider integration patterns.
- Drive model and architecture decisions using systematic evaluation of quality, latency, reliability, security, scalability, and cost.
- Define evaluation strategies, golden datasets, regression suites, observability, and operational SLOs for production AI systems.
- Design resilience patterns including fallbacks, caching, retries, human review, failure recovery, and safe degradation.
- Identify and build shared components that should be standardized across pods rather than repeatedly implemented.
- Lead technical design reviews, document architecture and trade-offs, and partner with Staff Engineers/Solution Advisory Board on cross-platform implications.
- Mentor AI Engineers I/II and raise engineering quality through code/design reviews and technical coaching.
- Evaluate emerging models, protocols, frameworks, and platform capabilities and recommend adoption when they create measurable value.
Technical Skills & Technology Stack
- Programming & Systems: Advanced Python and strong software/system design fundamentals; ability to integrate with Java or JavaScript/TypeScript services as needed.
- Data: Strong SQL, Snowflake or equivalent data platforms, data pipelines, structured/unstructured data, and retrieval-oriented data design.
- Generative AI / LLMs: Deep hands-on experience with multiple model providers, tool/function calling, structured generation, context engineering, and model-selection strategies.
- AI Architecture: Production RAG, vector/hybrid search, ranking, agents, tool orchestration, context/memory patterns, and human-in-the-loop workflows.
- Protocols & Frameworks: LangGraph/LangChain/LlamaIndex or equivalent; MCP/tool integration patterns; ability to design abstractions independent of individual frameworks/providers.
- Cloud & Platform: AWS or equivalent, containers, CI/CD, service observability, distributed system fundamentals, and production operations.
- Production AI: Evaluation frameworks, tracing/observability, guardrails, security, caching, fallbacks, latency/cost optimization, and failure-mode analysis.
Skills: langchain,ai,ci cd,rag,ml,python,mcp,llm,sql,aws,langgraph