Technical Lead

Crystal Peak
Noida, Uttar Pradesh, India

Role:Technical Lead-AI & Platform Services

Experience: 7+ Years

Location: Noida

Immediate Joiner

Role Overview:

We are looking for a Technical Lead- AI & Platform Services with strong hands-on software engineering experience to lead the design, development, deployment and production ownership of scalable backend and AI-enabled services.

The ideal candidate should have strong experience in Python, event-driven and queue-based architectures, pipeline/workflow orchestration, caching, cloud technologies and production reliability. Azure experience is preferred.

You will work closely with Software Engineers, AI Engineers, DevOps and AIOps teams and will co-own the reliability, performance and operational stability of services in production.

Key Responsibilities:

Software Engineering & Technical Leadership:

  • Lead the architecture, design and development of scalable backend services and APIs.
  • Remain hands-on with development, preferably using Python; Java or other enterprise languages are a plus.
  • Conduct code and design reviews and establish good engineering practices.
  • Mentor engineers and provide technical direction on architecture, performance, security and scalability.
  • Identify and address technical debt and continuously improve the platform.Event-Driven & Distributed Services:
  • Design and develop event-driven, queue-based and asynchronous services.
  • Work with technologies such as Azure Service Bus, Kafka, RabbitMQ, Event Grid or equivalent.
  • Design reliable mechanisms for retries, dead-letter queues, idempotency, timeouts and failure recovery.
  • Build services that can scale reliably with changing workloads.Pipeline & Workflow Orchestration:
  • Design and maintain business, data and AI processing pipelines.
  • Orchestrate workflows across APIs, services, queues and AI components.
  • Ensure workflows have appropriate retry, monitoring, failure recovery and scalability mechanisms.Caching & Performance:
  • Design and implement caching and in-memory data strategies to improve performance and scalability.
  • Experience with Redis or similar distributed caching technologies.
  • Address cache consistency, invalidation, TTL, memory usage and scalability considerations.AI / LLM Services:
  • Work closely with AI Engineers to integrate and productionize AI/LLM capabilities.
  • Build reliable backend services around AI models and APIs.
  • Support AI/LLM, RAG, agent and orchestration use cases.
  • Ensure AI services meet production requirements for reliability, scalability, observability and performance.
  • AI/LLM experience is preferred.Cloud & Deployment:
  • Design and deploy cloud-native services, with Azure preferred.
  • Experience with AKS, Docker, Kubernetes, Azure Service Bus, Azure Functions, Key Vault and Azure Monitor/Application Insights is valuable.
  • Work closely with DevOps on CI/CD, deployment, infrastructure, security and scalability.Production Reliability & Ownership:
  • Co-own production stability with DevOps/AIOps.
  • Take accountability for service reliability, performance and operational health.
  • Participate in incident resolution, RCA and permanent corrective actions.
  • Establish appropriate logging, monitoring, alerting, metrics, tracing and health checks.
  • Drive production readiness and post-release validation.
  • Proactively identify reliability, performance and scalability risks.Documentation & Release Management:
  • Own and maintain technical documentation for services and platforms.
  • Maintain architecture diagrams, API documentation, service dependencies, data flows, deployment procedures and operational SOPs.
  • Maintain release notes and ensure changes are properly documented.
  • Drive release readiness, deployment planning, rollback planning and post-release validation.Security & Engineering Quality:
  • Promote secure coding practices aligned with OWASP/SANS principles.
  • Ensure appropriate protection of credentials, secrets, PII and sensitive information.
  • Drive automated testing, code quality, dependency management and vulnerability remediation.Required Skills & Experience:
  • 7+ years of software engineering experience, preferably with significant backend/service development.
  • Strong hands-on Python experience; Java/C#/Go or similar languages are a plus.
  • Strong understanding of microservices, distributed systems, REST APIs, event-driven architecture and asynchronous processing.
  • Experience with messaging/queue technologies such as Azure Service Bus, Kafka or RabbitMQ.
  • Experience with Redis or equivalent caching technologies.
  • Strong cloud experience; Azure preferred.
  • Experience with Docker, Kubernetes/AKS and CI/CD.
  • Strong production troubleshooting, observability and reliability experience.
  • Good understanding of SQL/NoSQL databases and automated testing.
  • Strong technical leadership, communication and problem-solving skills.Preferred Experience:
  • Experience building AI/LLM-powered applications or services.
  • Experience working closely with AI/ML Engineers.
  • Experience with LLM APIs, RAG, AI Agents or AI orchestration.
  • Experience with Azure OpenAI or similar AI platforms.
  • Experience with Infrastructure as Code and modern observability platforms.

What We Look For:

A strong ownership mindset—someone who can take a service from:

Design → Build → Test → Deploy → Monitor → Operate → Document → Improve

and take accountability for its technical quality, reliability and production stability.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →