Research Engineer, Multi-Domain Alignment (SLM)

Invyte.ai
Gujarat, India

About Invyte, IncAn AI-native hiring platform advancing Sovereign AI through compact, "right-sized" models that are steerable, reliable, and deployable across high-stakes domains.

Job DescriptionBridge the gap between base pretraining and real-world deployment by architecting behavioral logic and safety frameworks for a new class of multimodal SLMs. Focus on multi-domain alignment to ensure models can transition seamlessly between specialized fields—Legal, Healthcare, and Industrial Robotics—while maintaining rigorous adherence to human intent and cultural values.

Key Responsibilities

  • Design and implement scalable alignment pipelines (SFT, DPO, PPO) to optimize 1B–7B parameter models for high-stakes, domain-specific tasks
  • Architect reward models and preference datasets that capture nuanced domain expertise, moving beyond generic helpfulness to expert-level reasoning
  • Develop innovative techniques to mitigate alignment drift and catastrophic forgetting when models are specialized across disparate industries
  • Devise rigorous, automated benchmarking suites (LLM-as-a-judge) and adversarial testing frameworks to validate model robustness in out-of-distribution scenarios
  • Contribute to the broader AI community by open-sourcing high-quality code and producing reproducible researchQualifications
  • Master's or PhD in Computer Science, ML, or equivalent practical experience in training large-scale models
  • Expertise in Python and PyTorch, specifically within the Hugging Face ecosystem (Transformers, TRL, PEFT, Accelerate)
  • Significant experience with RLHF, Direct Preference Optimization (DPO), and Constitutional AI
  • Deep understanding of Scaling Laws and the Alignment Tax—maximizing performance in compute-constrained environments
  • Experience aligning models that process text, visual, and sensor-based data
  • Research results published at leading venues such as NeurIPS, ICML, ICLR, or MLSys (bonus)
  • Experience building high-fidelity synthetic data pipelines to improve multi-step reasoning and logic (bonus)
  • Familiarity with optimizing inference engines (vLLM, TensorRT-LLM) or writing custom kernels (Triton/CUDA) for edge deployment (bonus)

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →