AI / ML & Generative AI Engineer
Title: AI / ML & Generative AI Engineer
Location: India Remote
Duration: 12+ month contract
Compensation: Competitive, based on experience
Work Requirements: Citizen of the country of employment or authorized to work in that country.
Senior AI and ML Operations Engineer
Serve as the primary technical resource for production generative AI and machine learning operations. This single individual-contributor position combines LLM gateway and retrieval support with model deployment, inference, evaluation and pipeline operations. Restore service, maintain output quality, improve monitoring and release controls, and automate recurring operational tasks across both areas.
The role requires demonstrated production experience in both generative AI and conventional machine learning, strong troubleshooting judgment and the ability to prioritize competing demands. Work closely with application, data, cloud and model owners to resolve incidents and deliver sustainable operational improvements.Technology environment
Python, SQL, Azure Foundry/Azure OpenAI and other LLM provider APIs, Azure AI Search, Databricks, Snowflake, AKS and Azure App Service. Supporting dependencies may include Redis, Couchbase, Oracle and Azure Blob. Git, CI/CD and observability tools support release and operational workflows. Application-specific gateways, model frameworks and configurations are learned during onboarding.Key responsibilities
- Provide Level 2/Level 3 support through ServiceNow. Assess business impact, diagnose incidents, coordinate recovery and escalation, and document root causes, workarounds and corrective actions.
- Troubleshoot LLM gateways and provider integrations, including authentication, quotas, rate limits, routing, timeouts, retries and approved fallback behavior. Trace failures across APIs and downstream services.
- Operate retrieval-augmented generation (RAG) and search workflows. Investigate indexing failures, stale knowledge, embeddings, filtering and relevance issues in Azure AI Search and connected data pipelines.
- Support Databricks model, processing and evaluation jobs and batch or online inference. Investigate runtime failures, dependencies, schema changes, data freshness, scoring errors and degraded predictions with data and model owners.
- Maintain evaluation datasets and automated checks for retrieval relevance, groundedness, response safety and model quality. Monitor drift and regressions using suitable baselines and business acceptance criteria; coordinate retraining or model changes with accountable owners.
- Monitor availability, latency, errors, throughput, token usage/cost, job success and output freshness. Correlate logs, metrics and traces to distinguish data, model, provider, application and infrastructure failures.
- Version code, prompts, configurations and model artifacts. Maintain reproducible environments, automated tests and controlled release, promotion and rollback procedures using Git and CI/CD.
- Support vision, image-processing or OCR workflows when assigned. Diagnose input and output-quality issues and engage domain specialists for deeper model or algorithm changes.
- Automate repeatable checks and recovery tasks. Maintain runbooks, dependency maps, handovers and escalation paths; apply approved access controls, secret handling, sensitive-data protection and AI safety practices. Required qualifications
- Demonstrated production support or engineering experience across both generative AI services and machine learning pipelines, including incident response, deployments, monitoring and recovery. Be able to explain direct contributions and operational outcomes in each area.
- Hands-on enterprise LLM API integration and gateway troubleshooting, plus practical RAG, embeddings, vector or hybrid search, retrieval filtering and response evaluation experience.
- Hands-on Databricks or comparable ML-platform experience, including job diagnostics, model/artifact versioning, reproducibility, batch or online inference, evaluation and rollback.
- Strong Python and SQL skills for debugging, data investigation and automation. Experience diagnosing REST APIs, authentication and service-to-service failures.
- Experience with Git, CI/CD, automated testing and controlled releases. Working knowledge of containers, Kubernetes application diagnostics and cloud service dependencies.
- Ability to select useful monitoring and evaluation measures, distinguish data problems from model regressions, and communicate business impact and technical findings clearly.
- Ability to prioritize a shared incident and improvement backlog, document operational procedures and train backup resources. A relevant degree or equivalent professional experience.Preferred experience Azure Foundry/Azure OpenAI, Azure AI Search, Databricks and Snowflake production experience are strongly preferred. Additional useful experience includes MLflow or another model registry, Redis, Couchbase, Oracle, Azure Blob, AKS, App Service, Azure DevOps, Terraform, distributed tracing, Model Context Protocol integrations, recommendation systems, computer vision and OCR. ServiceNow and ITIL-aligned incident, problem and change practices are beneficial. Certifications supplement demonstrated production capability.Scope and working relationships Own diagnosis and operational coordination for assigned AI/ML services. Partner with application engineers on APIs and code defects, data engineers on data pipelines and stores, and cloud engineers on infrastructure. Model and business owners define acceptance criteria and approve material behavior changes. Deep database/cluster administration, new-model research and major feature development remain with the accountable specialist teams.
Why Join INSPYR Global Solutions (IGS)
Become part of a diverse and multicultural workplace where you'll collaborate with talented professionals from around the world, build meaningful connections, and grow your career in a supportive environment. At IGS, we're committed to helping you succeed by offering:
- Flexible remote and hybrid work opportunities
- Continuous learning and professional development programs
- A culture that prioritizes employee wellbeing
INSPYR Global Solutions is the global delivery team for INSPYR Solutions, connecting top talent with opportunities across IT and professional services. Through collaboration, expertise, and innovation, we support global teams and help drive impactful solutions for our clients. Our focus is building a community of skilled, motivated, and reliable professionals who contribute to growth, performance, and long-term success. Learn more about us at www.inspyrglobalsolutions.com.
INSPYR Solutions provides Equal Employment Opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, or any other protected status. INSPYR Solutions complies with all applicable laws governing nondiscrimination in employment in every location in which the company has facilities. Applicants requiring reasonable accommodation during the application or interview process should contact HR@inspyrsolutions.com for assistance.
Information collected and processed through your application with INSPYR Solutions (including any job applications you choose to submit) is subject to INSPYR Solutions' Privacy Policy and INSPYR Solutions' AI and Automated Employment Decision Tool Policy: https://www.inspyrsolutions.com/policies/. By submitting an application, you are consenting to being contacted by INSPYR Solutions through phone, email, or text.