Data Scientist – LLM Fine-Tuning (On-Premises AI)

Telence Solutions
Hyderabad, Telangana, India

About Telence Solutions

Telence Solutions is a multi-vendor network engineering, cybersecurity and AI operations company serving service providers and mission-critical industries, including power utilities, oil and gas, transportation and public safety. Our in-house platform, the Autonomic Network Engine (ANE), is a next-generation AI-driven network management system. It detects faults, finds the root cause, checks the impact of a fix, waits for an engineer's approval, applies the fix and verifies that the network recovered.

ANE is designed so that AI explains the diagnosis rather than guessing it. A causal graph determines the root cause, and language models explain findings, assist engineers and draw on each customer's own runbooks. Because many of our customers run sensitive networks, ANE runs fully on-premises, including air-gapped, or in AWS, with the language model fully swappable.

The Role

We're looking for a Data Scientist to own how the language models inside ANE are adapted to the realities of network operations. You'll turn real operational data (alarms, syslogs, telemetry, configurations, runbooks and engineers' decisions) into training and evaluation datasets. You'll then fine-tune open-weight models that run entirely on customer premises. Your work directly improves how accurately ANE explains faults, how well it follows a customer's playbooks, and how reliably it supports engineers.

What You'll Do

  • Curate operational data: build datasets from network alarms, syslogs, SNMP and streaming telemetry, configuration data, incident history and runbooks, including cleaning, deduplication, labeling and anonymization of customer data.
  • Turn engineer decisions into training signal: ANE records every proposed fix and whether an engineer approved, edited or rejected it. Use that record to build supervised and preference datasets.
  • Fine-tune open-weight models: retrain and fine-tune LLMs and smaller language models (for example Llama, Mistral, Qwen or DeepSeek) using LoRA/QLoRA, supervised fine-tuning and preference-optimization methods, sized for customer GPU hardware.
  • Build evaluation harnesses: measure domain accuracy, hallucination rates, tool-calling reliability and faithfulness to runbooks, so that no model ships without passing defined quality gates.
  • Improve retrieval: fine-tune embedding models and tune RAG pipelines over vendor documentation and customer playbooks.
  • Prepare models for on-premises deployment: handle quantization, packaging and versioning so models run efficiently with local inference servers (for example vLLM) in air-gapped environments.
  • Support forecasting: contribute time-series models for capacity and bandwidth forecasting and device-health trends.
  • Run the model lifecycle: maintain reproducible training pipelines, model registries (for example MLflow), model cards and governance documentation.
  • Work across the team: collaborate with our AI architect, software developers and network engineers, who carry more than 30 years of experience building carrier and enterprise networks.
  • What You Bring
  • 4+ years in data science or machine learning, including 2+ years of hands-on LLM fine-tuning.
  • Strong Python and PyTorch, plus experience with Hugging Face Transformers, PEFT and TRL.
  • Practical experience building fine-tuning datasets, including data cleaning, labeling strategy and synthetic data generation.
  • Experience designing LLM evaluations and benchmarks, not just training runs.
  • Comfort working on Linux GPU servers, with distributed training (DeepSpeed or FSDP) and model quantization.
  • Experience with RAG systems, embedding models and vector databases.
  • Clear communication, and the judgment to know when a model isn't ready to ship.
  • Nice to Have
  • Experience with network or IT operations data: syslog, SNMP, streaming telemetry, alarms or trouble tickets.
  • Familiarity with IP/MPLS, optical or SD-WAN networks, or with network management platforms such as Nokia NSP, Cisco Crosswork, Juniper Paragon, Grafana or Prometheus.
  • Experience deploying models in air-gapped or highly regulated environments.
  • Time-series forecasting and anomaly detection experience.
  • Knowledge of agentic AI, multi-agent systems or knowledge graphs.
  • Why Join Us
  • Shape the AI at the core of a product built for mission-critical networks, where accuracy matters more than demos.
  • Work with real operational data and real engineers, not toy benchmarks.
  • See your models running in production on customer premises, under a design that keeps a human in control of every change.
  • Join a small, senior team where your decisions directly shape the platform.If you want to make language models genuinely reliable in a domain where mistakes cause outages, we'd love to hear from you.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →