Principal Data Scientist

Optum India
Bengaluru, Karnataka, India

Key Responsibilities

  • Deep knowledge and extensive experience with Machine/Deep Learning frameworks including transformer architectures, state space models, large language models, and agentic approaches
  • Lead end-to-end development and implementation of Large Language Model (LLM) solutions across both open-source (e.g., Qwen, LLaMA, Mistral) and closed-source (e.g., OpenAI, Gemini, Anthropic) ecosystems.
  • • Knowledge of algorithms and techniques within a computational domain with emphasis on text processing
  • Demonstrated publication record in AI domain especially relating to text extraction and summarization
  • Architect and implement GraphRAG pipelines, including knowledge graph representation and retrieval for enhanced contextual grounding.
  • Design, train, and optimize semantic and dense vector embeddings for document understanding, search, and retrieval.
  • Develop semantic retrieval systems with advanced document segmentation and indexing strategies.
  • Build and scale distributed training environments using NCCL and InfiniBand for multi-GPU and multi-node training.
  • Apply reinforcement learning techniques (e.g., RLHF, RLAIF) to align model behavior with human preferences and domain-specific goals.
  • Experience with Hybrid NLP solutions that combine symbolic and machine learning approaches
  • Collaborate with cross-functional teams to translate business needs into AI-driven solutions and deploy them in production environments.

Required Qualifications

Master’s degree in computer science, Machine Learning, or related field.

11-15+ years of experience in applied AI/ML with statistics, with a strong track record of delivering production-grade models.

Deep expertise in: NLP, Fundamental machine learning, deep learning, transformer, state space-based architecture

Azure ML and/or AWS

Strong in Python coding, SQL and database queries, data preparation, and analysis

Exploratory Data Analysis (EDA)

Experience with PyTorch

Fine-tuning (e.g., GPT, LLaMA, Mistral, Qwen)

Graph-based retrieval systems (knowledge graphs)

Embedding models (e.g., BGE, E5, SimCSE)

Semantic search and vector databases (e.g., FAISS, Weaviate, Milvus)

Model fusion and ensemble techniques (stacking, boosting, gating)

Optimization algorithms (Bayesian, Particle Swarm, Genetic Algorithms)

Reinforcement learning (e.g., RLHF, PPO, DPO, GRPO), Supervised Fine Tuning (SFT), LoRA, QLoRA, axolotl

Prompt optimization framework (AutoPrompt, GreaterPrompt, DSPy), GEPA

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →