Technical Lead
Role:Technical Lead-AI & Platform Services
Experience: 7+ Years
Location: Noida
Immediate Joiner
Role Overview:
We are looking for a Technical Lead- AI & Platform Services with strong hands-on software engineering experience to lead the design, development, deployment and production ownership of scalable backend and AI-enabled services.
The ideal candidate should have strong experience in Python, event-driven and queue-based architectures, pipeline/workflow orchestration, caching, cloud technologies and production reliability. Azure experience is preferred.
You will work closely with Software Engineers, AI Engineers, DevOps and AIOps teams and will co-own the reliability, performance and operational stability of services in production.
Key Responsibilities:
Software Engineering & Technical Leadership:
- Lead the architecture, design and development of scalable backend services and APIs.
- Remain hands-on with development, preferably using Python; Java or other enterprise languages are a plus.
- Conduct code and design reviews and establish good engineering practices.
- Mentor engineers and provide technical direction on architecture, performance, security and scalability.
- Identify and address technical debt and continuously improve the platform.Event-Driven & Distributed Services:
- Design and develop event-driven, queue-based and asynchronous services.
- Work with technologies such as Azure Service Bus, Kafka, RabbitMQ, Event Grid or equivalent.
- Design reliable mechanisms for retries, dead-letter queues, idempotency, timeouts and failure recovery.
- Build services that can scale reliably with changing workloads.Pipeline & Workflow Orchestration:
- Design and maintain business, data and AI processing pipelines.
- Orchestrate workflows across APIs, services, queues and AI components.
- Ensure workflows have appropriate retry, monitoring, failure recovery and scalability mechanisms.Caching & Performance:
- Design and implement caching and in-memory data strategies to improve performance and scalability.
- Experience with Redis or similar distributed caching technologies.
- Address cache consistency, invalidation, TTL, memory usage and scalability considerations.AI / LLM Services:
- Work closely with AI Engineers to integrate and productionize AI/LLM capabilities.
- Build reliable backend services around AI models and APIs.
- Support AI/LLM, RAG, agent and orchestration use cases.
- Ensure AI services meet production requirements for reliability, scalability, observability and performance.
- AI/LLM experience is preferred.Cloud & Deployment:
- Design and deploy cloud-native services, with Azure preferred.
- Experience with AKS, Docker, Kubernetes, Azure Service Bus, Azure Functions, Key Vault and Azure Monitor/Application Insights is valuable.
- Work closely with DevOps on CI/CD, deployment, infrastructure, security and scalability.Production Reliability & Ownership:
- Co-own production stability with DevOps/AIOps.
- Take accountability for service reliability, performance and operational health.
- Participate in incident resolution, RCA and permanent corrective actions.
- Establish appropriate logging, monitoring, alerting, metrics, tracing and health checks.
- Drive production readiness and post-release validation.
- Proactively identify reliability, performance and scalability risks.Documentation & Release Management:
- Own and maintain technical documentation for services and platforms.
- Maintain architecture diagrams, API documentation, service dependencies, data flows, deployment procedures and operational SOPs.
- Maintain release notes and ensure changes are properly documented.
- Drive release readiness, deployment planning, rollback planning and post-release validation.Security & Engineering Quality:
- Promote secure coding practices aligned with OWASP/SANS principles.
- Ensure appropriate protection of credentials, secrets, PII and sensitive information.
- Drive automated testing, code quality, dependency management and vulnerability remediation.Required Skills & Experience:
- 7+ years of software engineering experience, preferably with significant backend/service development.
- Strong hands-on Python experience; Java/C#/Go or similar languages are a plus.
- Strong understanding of microservices, distributed systems, REST APIs, event-driven architecture and asynchronous processing.
- Experience with messaging/queue technologies such as Azure Service Bus, Kafka or RabbitMQ.
- Experience with Redis or equivalent caching technologies.
- Strong cloud experience; Azure preferred.
- Experience with Docker, Kubernetes/AKS and CI/CD.
- Strong production troubleshooting, observability and reliability experience.
- Good understanding of SQL/NoSQL databases and automated testing.
- Strong technical leadership, communication and problem-solving skills.Preferred Experience:
- Experience building AI/LLM-powered applications or services.
- Experience working closely with AI/ML Engineers.
- Experience with LLM APIs, RAG, AI Agents or AI orchestration.
- Experience with Azure OpenAI or similar AI platforms.
- Experience with Infrastructure as Code and modern observability platforms.
What We Look For:
A strong ownership mindset—someone who can take a service from:
Design → Build → Test → Deploy → Monitor → Operate → Document → Improve
and take accountability for its technical quality, reliability and production stability.