LLMOps/MLOps Lead

Klimber Technologies
India

Company Description Klimber Technologies is an AI-first company built by a team with decades of experience in enterprise data, analytics, and cloud platforms across AWS, GCP, and Azure. The organization focuses on building AI products that remove friction for enterprises, solving real problems and delivering production-grade solutions for Fortune 100 clients and beyond. Klimber’s vision is to enable businesses to move fast with AI while maintaining reliability and trust in the outputs, ensuring pilots have a clear path to production. Its offerings include Turn3, an agent monitoring solution that evaluates whether AI agents truly achieve user goals, and hypernear, a vector database and similarity search engine for large-scale RAG, semantic search, and agent memory. In addition to products, Klimber partners with enterprises and ISVs to deliver end-to-end LLMOps, domain-tuned language models, knowledge-graph-driven analytics, and human-in-the-loop moderation for regulated and high-stakes use cases.

Role Description The LLMOps/MLOps Lead will oversee the design, implementation, and operation of production-scale AI and LLM pipelines across Klimber products and client environments. This contract, remote role involves architecting and maintaining CI/CD workflows for machine learning models, deploying and monitoring LLM-based systems, and ensuring observability, reliability, and security of AI infrastructure. The lead will collaborate closely with product, engineering, and data science teams to define deployment standards, manage model lifecycle (training, evaluation, rollout, rollback), and optimize performance in multi-cloud and on-premise settings. Day-to-day work includes building tooling for agent monitoring, integrating vector databases and RAG pipelines, implementing guardrails and human-in-the-loop workflows, and translating operational insights into product and platform improvements. The person in this role will also mentor engineers on MLOps/LLMOps best practices and contribute to internal frameworks that help clients move AI solutions from pilot to production.

Qualifications

  • Strong experience with MLOps/LLMOps practices, including model lifecycle management, CI/CD for ML, and production monitoring of AI systems.
  • Hands-on proficiency with cloud platforms (AWS, GCP, Azure) and containerization/orchestration tools such as Docker and Kubernetes.
  • Solid background in Python-based ML/LLM development, including integration with frameworks such as PyTorch, TensorFlow, or popular LLM orchestration libraries.
  • Experience designing and operating retrieval-augmented generation (RAG) pipelines, vector databases, and semantic search or similarity search infrastructure.
  • Knowledge of observability and reliability tooling (e.g., logging, metrics, tracing, alerting) for AI and microservices, with a focus on safety, guardrails, and agent monitoring.

-

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →