Senior Voice AI / Speech ML Engineer

TalixoHR
Greater Bengaluru Area

Experience: 3–4 years

Employment: Full-time

Location: HSR Layout Bengaluru | On-site

About the Role

We are looking for a Senior Voice AI / Speech ML Engineer to join our AI team and build the next generation of speech and voice intelligence systems.

This role is for an engineer/researcher who has hands-on experience training large-scale speech or language foundation models, not just integrating or fine-tuning existing models.

You will work across ASR, TTS, speaker identification, VAD and conversational AI, taking models from large-scale data preparation and training through evaluation, optimization and production deployment.

If you have personally trained/pre-trained a foundation model and have strong ASR/TTS experience, we would like to hear from you.

What You'll Work On

  • Design, train, pre-train and fine-tune foundation models for Voice AI.
  • Develop and train ASR and/or TTS models for real-world speech applications.
  • Work with large-scale speech, audio and text datasets.
  • Build data preparation, preprocessing, augmentation and model-training pipelines.
  • Design experiments and evaluate models across accuracy, robustness and generalization.
  • Improve WER/CER, speech quality, latency and inference efficiency.
  • Work with Transformer-based architectures and modern speech-model architectures.
  • Train models using GPUs and distributed computing infrastructure.
  • Debug training instability, data issues and model-performance bottlenecks.
  • Optimize trained models for production, including quantization and inference optimization.
  • Collaborate with research and engineering teams to move models from experimentation to production.

Must-Have Experience

  • 3–4 years of hands-on experience in Speech ML, Speech AI, NLP, Deep Learning or Machine Learning.
  • Actual foundation-model training/pre-training experience.
  • Hands-on experience developing or training ASR and/or TTS models.
  • Strong understanding of Transformers and deep learning.
  • Strong Python programming skills.
  • Hands-on experience with PyTorch and/or TensorFlow.
  • Experience with GPU-based model training and large datasets.
  • Strong understanding of model evaluation, experimentation and optimization.

Strongly Preferred

Experience with one or more of:

  • Whisper
  • wav2vec 2.0
  • HuBERT
  • Conformer / FastConformer
  • NVIDIA NeMo
  • SpeechBrain
  • Hugging Face
  • F5-TTS / VITS / similar TTS architectures
  • Speech or audio foundation models
  • Multilingual / low-resource speech
  • Indian-language speech datasets
  • Distributed training using DeepSpeed, FSDP, Ray Train or similar
  • Audio preprocessing, augmentation, annotation and dataset-quality pipelines
  • Production deployment of speech models

What We Are Specifically Looking For

The ideal candidate has worked on the model itself, rather than only building applications around existing models.

Strong fit:

Foundation-model pretraining + ASR/TTS + Transformers + PyTorch + GPU/distributed training + large-scale speech dataNot the right fit:

  • Generic ML Engineer without speech experience
  • GenAI / RAG Engineer
  • Prompt Engineer
  • LLM application developer
  • Candidates who only consume OpenAI/Claude/Gemini APIs
  • Candidates who only fine-tune existing pretrained models
  • Conversational AI developers who integrate speech APIs/Riva/voice APIs but don't train speech models
  • Candidates with only NLP/LLM experience and no ASR/TTS

Why Join

  • Work on core Voice AI / Speech ML technology, not only application-layer GenAI.
  • Solve challenging problems involving large-scale model training, speech data and inference optimization.
  • Work closely with AI/ML engineers and researchers on production-grade models.
  • Opportunity to work on multilingual and real-world speech applications in a high-growth AI environment.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →