Inference Engineer

Antrino Labs
Greater Bengaluru Area

We are looking for an excellent AI Backend Engineer to join one of Sweden's fastest growing start-ups which is backed by Nvidia, Microsoft and AWS. The company also operates under the Inspection for Strategic Products.

Location:

Fully remote within India. From 2027 you join our Bangalore office as part of the founding local team Employment: Full-time, permanent Start: As soon as you are available

About the job:

Antrino Labs is building the platform that makes every square meter intelligent.

The physical world runs on processes nobody can actually see. Goods move, people queue, machines idle, space goes unused, and the decisions made about all of it are based on samples, guesses, and reports written after the fact. Software solved this for the digital world twenty years ago. Everything online is measured, understood, and acted on in real time. The physical world is still dark.

We are building the execution layer that closes that gap. Antrino lets any organization deploy vision intelligence into a physical space and get back what is actually happening there — as structured, queryable, real-time information rather than footage. Not a research project, not a custom integration, not a team of ML engineers. A platform, where you describe what matters to you and deploy it.

What people build on it is broader than what we designed for. Operators use it to understand processes, find where time and space are being wasted, measure flow and utilization, catch problems while they are still happening, and turn all of it into data good enough to act on.

Everything is built on Privacy by Design. Data protection and GDPR compliance are part of the architecture, not a layer added afterwards.

We recently closed our Seed round, announced together with Dagens Industri, and we are scaling to meet demand. Antrino Labs is also registered with Inspektionen för strategiska produkter (ISP), the Swedish authority for strategic and dual-use products.

The role:

We are hiring an experienced AI Backend Engineer to build the services that turn a customer's intent into working intelligence.

This is distinct from our MLOps and DevOps roles, and it is worth being precise about the difference. MLOps owns whether a model runs well in production, and DevOps owns the infrastructure everything runs on. You own the layer in between: the backend logic that lets a customer describe what they want detected, compiles that into a real execution plan, routes it across our detection backends and model APIs, and turns raw model output into something structured, explainable, and actionable. If MLOps is about models running fast and DevOps is about the ground they run on, you are building the reasoning and orchestration that sits on top of both.

Concretely, when a customer builds something in Studio, our composition canvas, that graph has to become a running system: a plan that decides which models to call, in what order, gated by cost and confidence, producing events and evidence that a customer can trust. You will work on the compiler that turns a graph into an execution plan, the runtime that executes it, and the backend services — written in Python and TypeScript — that detection, classification, and measurement actually run through. You will also work on our Agent, which interviews customers, proposes and edits these graphs, and deploys them with confirmation. Getting an LLM to reliably manipulate a structured graph, validate it, and explain its own reasoning is a real engineering problem, not a prompting exercise.

You will work closely with the team building our first-party detection models and with the product engineers building the interfaces customers use, but this role sits between them: it is backend systems engineering for a product whose core logic is built out of AI calls.

Qualifications

  • Strong programming skills in languages commonly used for ML and systems engineering (e.g., Python) and experience writing clean, maintainable, testable code.
  • Hands-on experience with machine learning model deployment and inference frameworks (e.g., TensorRT, ONNX Runtime, or equivalent tools).
  • Solid understanding of distributed systems, microservices, and API-based architectures, Fast API services.
  • Knowledge of performance optimization techniques, including profiling, hardware acceleration (GPU), and parallelization for low-latency inference.
  • Experience working with containerization and orchestration technologies such as Docker for production deployments.
  • Familiarity with cloud platforms (AWS) and related services for scalable model hosting and monitoring.
  • Proficiency in using CI/CD pipelines to support reliable, repeatable releases.
  • Strong analytical and problem-solving abilities, with attention to reliability, robustness, and security of inference systems.
  • Effective communication skills and the ability to collaborate with cross-functional teams in an on-site environment.
  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, Data Science, or a related technical field, or equivalent practical experience.
  • Experience with logging, monitoring, and observability tools (e.g., Prometheus, Grafana, or similar) to track performance and system health.
  • Background in applied machine learning or MLOps is a plus, especially in production settings.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →