Senior Voice AI / Speech ML Engineer | Foundation Models | 3–4 Years

TalixoHR
Greater Bengaluru Area

Build the intelligence behind next-generation Voice AI.

A high-growth AI technology team is looking for a hands-on Voice AI / Speech ML Engineer who has already worked on training and fine-tuning foundation models.

This is not a generic ML role. You’ll work directly on speech models, large-scale datasets, model training and production optimization across ASR, TTS and conversational AI.

ABOUT THE OPPORTUNITY

You’ll own model development from data preparation and experimentation through training, evaluation and deployment, with a focus on improving accuracy, latency, robustness and inference efficiency.

The role is ideal for an ML engineer who wants deeper ownership of foundation-model training and real-world Voice AI systems.

WHAT YOU'LL OWN

  • Develop, train, fine-tune and evaluate foundation models for Voice AI.
  • Build solutions across ASR, TTS, speaker identification and voice activity detection.
  • Work on conversational AI and speech-language models.
  • Prepare and optimize large-scale speech and text datasets.
  • Design model-training pipelines and run experiments.
  • Improve model accuracy, latency, robustness and inference efficiency.
  • Work with distributed training and GPU-based workloads.
  • Optimize models for production using techniques such as quantization.
  • Debug model and pipeline issues and drive continuous performance improvements.
  • Support production deployment of trained models.

THE IDEAL CANDIDATE

We are looking for someone who has actually trained or pre-trained foundation models, not someone whose experience is limited to consuming existing APIs or models.

MUST HAVE

  • 3–4 years of hands-on experience in ML, Speech AI, NLP or related fields.
  • Proven experience training/pre-training foundation models.
  • Strong understanding of Transformers and deep learning architectures.
  • Hands-on experience with ASR and/or TTS models.
  • Strong Python skills.
  • Hands-on PyTorch and/or TensorFlow.
  • Experience with distributed training, GPUs and large-scale data pipelines.
  • Experience with model evaluation and optimization.
  • Strong debugging and problem-solving skills.

GOOD TO HAVE

  • Multilingual or Indian-language speech models.
  • Whisper, wav2vec 2.0, HuBERT, NeMo, SpeechBrain or Hugging Face.
  • Audio preprocessing, augmentation and annotation.
  • Dataset quality improvement.
  • Production model deployment.

WHY THIS ROLE

  • Work directly on Voice AI and foundation-model technology.
  • Go beyond API integration into actual model training and optimization.
  • Work with large-scale speech datasets and GPU infrastructure.
  • Solve challenging problems across accuracy, latency and inference efficiency.
  • Build technology with direct impact on next-generation conversational experiences.

ROLE DETAILS

Role: Senior Voice AI / Speech ML Engineer

Experience: 3–4 Years

Domain: Voice AI / Speech ML / NLP

Employment: Full-Time

Joining: As per business requirement

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →