AI Engineer (LLM & Agent Systems)

Speegile Consulting
Mumbai, Maharashtra, India

About the RoleWe are looking for an experienced AI Engineer to design, build, and deploy LLM-driven agentic systems end-to-end — from data and model fine-tuning to retrieval, evaluation, guardrails, and production deployment. You should be equally comfortable reasoning about how a transformer works internally and shipping a reliable, observable agent system in production. This role owns the AI/ML depth of our stack, with support from backend and frontend engineers for integration.

Experience: 3+ Years

Qualification: B.E. / B.Tech / M.Sc. / M.Tech / MCA — Computer Science / IT or related field

Location: Preference Mumbai / Relocation to Mumbai / Remote

Employment Type: Full-Time

Notice Period: Max 30 Days

Must-Have Skills• 3+ years of experience building Python-based ML/AI systems in production

  • Strong, practical understanding of how transformer, encoder-decoder, and deep learning models work internally — not just API-level usage
  • Hands-on experience fine-tuning LLMs (LoRA / QLoRA / full fine-tuning) and working with large-scale datasets end-to-end
  • Foundational Knowledge of LLMs and Can Lay the foundation for building a LLM model from scratch.
  • Solid grasp of the full LLM landscape — RAG, vectored and vector-less retrieval, evaluation, observability, and guardrails
  • Experience with voice/speech models (ASR/TTS & Translation) in Fine Tuning the models can even lay the foundation to build a voice model from scratch
  • Knowledge graphs, Redis, Advanced RAG , Self corrective RAG
  • Experience with cloud AI deployments (AWS / GCP / Azure)
  • Hands-on experience with LLM agent frameworks and vector databases (ChromaDB, Weaviate, pgvector)
  • Strong knowledge of PyTorch or TensorFlow and Scikit-Learn
  • Production experience with FastAPI, Docker, and MLOps

Good-to-Have Skills• Advanced prompt optimization and agent evaluation techniques

  • Experience building AI observability tools from scratch
  • Exposure to multilingual or low-resource language models
  • Expert-level usage of agentic coding IDEs (Cursor, Windsurf, Claude Code)

Soft Skills• Strong problem-solving and system design mindset

  • Ability to work across AI, backend, and front-end teams
  • Comfortable mentoring freshers/junior engineers on AI fundamentals
  • Clear communication and documentation skills
  • Passion for building production-grade AI systems

What You'll Work OnLLM Agents & Prompt Engineering

  • Design and implement LLM agents using LangGraph, PydanticAI, and Google ADK
  • Build tool-augmented reasoning pipelines: RAG, Chain-of-Thought, ReAct, and planner–executor architectures
  • Develop robust, tested prompt strategies to improve reliability and reduce hallucination

Model Fundamentals & Fine-Tuning

  • Deep working knowledge of transformer architecture, Mamba architecture— encoder-only, decoder-only, and encoder–decoder models — and how attention, embeddings, and positional encoding actually work under the hood
  • Fine-tune LLMs using LoRA, QLoRA, and full fine-tuning depending on the use case and compute budget. If the fine tuning doesn’t get the desired results, a LLM model will be built from scratch
  • Prepare, clean, and process large-scale datasets for pre-training/fine-tuning — deduplication, tokenization, sampling, and quality filtering at scale
  • Understanding of core deep learning model families (CNNs, RNNs/LSTMs, transformers) and when to use each
  • Work with voice/speech models — ASR (speech-to-text) and TTS (text-to-speech) — and understand how they integrate into conversational AI pipelines

Retrieval, RAG & Knowledge Systems

  • Implement both vectored (embedding-based) and vector-less (keyword/graph/hybrid) retrieval strategies, choosing the right approach per use case
  • Integrate vector databases (FAISS, Pinecone, pgvector, ChromaDB, Weaviate) and knowledge graphs (Neo4j)
  • Design chunking, embedding, and re-ranking strategies for high-precision retrieval
  • Implement Prompt Engineering, Context Engineering, Loop Engineering & Efficient Low token retrieval

Evaluation, Observability & Guardrails

  • Build offline and online evaluation harnesses for agent and LLM outputs (accuracy, groundedness, latency, cost)
  • Implement guardrails for safety, PII redaction, and scope control (e.g. NeMo Guardrails, Guardrails AI, or custom rule/LLM-based filters)
  • Build tooling for trace analysis, state debugging, and hallucination detection
  • Set up observability dashboards for LLM/agent systems (e.g. LangSmith, Arize, custom logging pipelines)
  • Benchmark agent orchestration frameworks for performance, cost, and reliability

Backend & MCP Integration

  • Build scalable APIs using FastAPI (sync & async execution)
  • Implement Model Context Protocol (MCP) for secure tool and data access
  • Manage agent state, context routing, and plugin-based workflows

MLOps & Deployment

  • Deploy and monitor models in cloud environments (AWS / GCP / Azure)
  • Work with model serving frameworks (e.g. vLLM, TGI) and apply quantization for efficient inference
  • Implement logging, observability dashboards, and automated recovery workflows

Front-End Collaboration

  • Build or collaborate on UI using React, TypeScript, or Next.js
  • Create seamless UI–API bridges for agent interactions and basic dashboards

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →