AI Engineer, Exp: 4-7 Yrs, Noida, WFO

Ramp Infotech
Noida, Uttar Pradesh, India

Job Description

AI Engineer (LLM / RAG / Document Intelligence)

Role Summary

We are seeking an AI Engineer to build the AI core of a new B2B platform being delivered to a client. Client and project details are confidential and will be disclosed at onboarding under a confidentiality agreement.

The systems operate in a domain where an incorrectly extracted value is a serious defect: accuracy, grounding and controllability take priority over raw capability. The role covers production LLM engineering end-to-end: retrieval-augmented generation over a restricted document corpus with strict source boundaries, document and PDF data-extraction pipelines that normalize inconsistent real-world specifications, NLP classification pipelines, semantic search, and conversational intake that converts informal user language into precise structured data. The engineer works within a small senior delivery team - a Solution Architect who owns the technical design, and a Senior Full-Stack Developer who consumes the engineer's APIs - and demonstrates completed work in fortnightly sprint reviews with client stakeholders present.

EolasFlow is an AI-native engineering team: AI-assisted development (Claude Code, Cursor, GitHub Copilot or equivalent) is the standard working method, and candidates are expected to already work this way.

Key Responsibilities

  • Design and build a production RAG system: chunking and embedding strategy, vector store, retrieval evaluation, citation-grounded answering, and strict source-boundary enforcement with refusal on out-of-bound queries
  • Build document-intelligence pipelines: PDF and table extraction from inconsistent source documents, unit and format normalization, deduplication, and human-audit workflows
  • Build NLP pipelines for content classification (signal vs noise), entity extraction and enrichment, and automated draft generation matched to a defined editorial voice
  • Build semantic search mapping natural-language intent to structured capability data
  • Build LLM-guided conversational intake converting informal language into precise structured specifications
  • Establish evaluation discipline: evaluation sets and regression harnesses ahead of tuning, evaluations running in CI, quantified quality reporting
  • Monitor and optimize cost, latency and quality across all LLM usage; make provider and model trade-offs explicit
  • Expose all capabilities as clean, documented APIs for consumption by the application layer
  • Present completed work in fortnightly sprint reviews
  • Optimize high-volume pipeline tasks (classification, drafting) by fine-tuning and deploying small open-weight language models where they outperform API models on cost and latency

Required Skills and Experience

  • Minimum 4 years building ML/NLP/LLM systems in production; strong Python (FastAPI or similar for serving)
  • Production RAG experience: candidates must be able to walk through a shipped system — architecture, evaluation results, failure modes and remediation
  • LLM engineering: prompt design, structured output (JSON schema / function calling), multi-provider model selection (OpenAI, Anthropic, open-weight models), cost and latency optimization
  • Vector stores (pgvector, Qdrant, Pinecone or Weaviate); retrieval evaluation and hallucination control
  • Document intelligence: PDF and table extraction from inconsistent real-world documents (Unstructured, Textract, Docling or custom pipelines)
  • Evaluation discipline: builds evaluation sets and regression harnesses as standard practice and can quantify quality improvements
  • Classic NLP fundamentals beyond prompting classification, named-entity recognition, entity resolution
  • Data pipeline orchestration (Airflow, Prefect or similar); compliant API and web data ingestion (rate limiting, terms-of-service awareness)
  • Optimize high-volume pipeline tasks (classification, drafting) by fine-tuning and deploying small open-weight language models where they outperform API models on cost and latency
  • Experience building agentic pipelines (tool use, multi-step agents) in production
  • Daily, fluent use of AI-assisted development tools (Claude Code, Cursor, GitHub Copilot or equivalent); this will be assessed through a live practical exercise during selection
  • Fluent written and spoken English; able to present work to non-technical stakeholders

Desirable

  • Small language model (SLM) fine-tuning: LoRA/QLoRA adaptation of open-weight models (Llama, Mistral, Phi or similar class) for classification and style/domain adaptation, including serving and deployment (vLLM, Ollama or similar) valued as a cost- and latency-optimization path for high-volume pipeline tasks
  • Hybrid retrieval and re-ranking (BM25 combined with dense retrieval)
  • Experience with technical or industrial specification data
  • Content personalization or recommender systems

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →