Artificial Intelligence Engineer
Astra Security
Bengaluru, Karnataka, India
AI Engineer, Autonomous Pentesting
Why this role exists
A frontier model on its own is not a pentester. What turns a model into a capability you can run against a customer's app every day, and at scale, is the harness around it: how you feed it the target, scaffold its reasoning, give it tools, orchestrate cheap models against expensive ones, score what it produces, and stop it from shipping garbage. That harness is our core IP and our main lever on accuracy, coverage, and cost. You will own it.
What you'll do
- Build and productionize harnesses for autonomous pentesting across web, API, mobile, cloud, and more, taking each from a working prototype to something reliable against real customer targets.
- Orchestrate models against each other: use frontier models as the source of intelligence, then distill that into cheaper open-weight models so we can run at scale economically.
- Own evaluation. Define what a true finding is, measure recall and precision against ground truth, penalize false positives hard, normalize for cost, and control for run-to-run variance. If it is not measured, it does not ship.
- Push accuracy and coverage up, and cost and latency down, release over release.
- Build the guardrails that keep a hallucinated finding from ever reaching a customer, because one false critical costs more trust than ten real findings earn.
- Work shoulder to shoulder with our pentesters and researchers: turn their judgment into harness logic, and build tools that make them faster.What we're looking for
- An engineer who ships. You write clean, scalable, flexible code, get a working slice out fast, then harden it. Prototypes that never reach production are not the job.
- Real LLM engineering: prompting and scaffolding, tool use and agents, orchestration across models, retrieval and context management, and a feel for where models break.
- Enough security depth to judge the output. You do not need to be a career pentester, but you must look at a finding and tell real from plausible-but-wrong, and reason about severity. Slop resistance is a core skill here.
- Evaluation rigor. You think in ground truth, baselines, false-positive rates, and cost per true finding, not vibes.
- Production instincts: observability, cost control, reliability, and the discipline to measure before and after every change.Nice to have
- Hands-on pentesting or offensive security across web, API, mobile, or cloud, CTF, or vulnerability research.
- Serving or fine-tuning open-weight models, or running them cheaply at scale.
- Mobile (Android or iOS) security, static and dynamic analysis, or reverse engineering.
- A track record of turning a messy research idea into something that shipped.