Infrastructure System Engineer
About Nava
Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the metal, on infrastructure built to keep GPUs saturated and models serving.
About The Role
We are building a next-generation Infrastructure Observability Platform that provides a unified view of modern datacenter and cloud infrastructure across physical, network, compute, GPU, storage, and software layers. The platform will build a rich infrastructure graph linking assets, dependencies, workloads, tenants, and services, enabling real-time visibility into health, capacity, performance, topology, and utilization. The role involves designing and building scalable data pipelines, asset and topology models, telemetry ingestion, correlation, and intelligent analytics for large-scale infrastructure environments. You will work across distributed systems, cloud infrastructure, Kubernetes, GPUs, networking, and data platforms to create a platform that turns infrastructure data into actionable operational intelligence.
What You’ll Do
- Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.
- Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation.
- Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers.
- Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators.
- Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.
- Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations.
- Continuously improve deployment velocity, reliability, and operational efficiency through automation.
- Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure.
What You’ll Need
- 4-7 years building distributed systems, infrastructure platforms, or large-scale backend software.
- Strong software engineering skills in Go, Python.
- Experience building platforms, automation systems, or developer infrastructure.
- Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies.
- Strong systems thinking with the ability to understand problems across hardware and software.
- A passion for solving complex infrastructure challenges through software.
- An automation-first mindset—if a task is repeated, your instinct is to build a system to eliminate it.
Skills: go (golang),cloud,kubernetes,ai infrastructure,gpu