AI/ML Engineer — Data, Evaluation & Model Improvement
AI/ML Engineer — Data, Evaluation & Model Improvement
Focus: LLM evaluation, fine-tuning and continuous product improvement
Company: SecNinjaz Technologies LLP
Experience: Around 2 years of relevant hands-on experience
Employment: Full-time, on-site
Locations: Netaji Subhash Place, New Delhi,
About the opportunity
At SecNinjaz, we are building an AI-powered Vulnerability Assessment and Penetration Testing (VAPT) product that helps organisations identify, investigate and validate security weaknesses in authorised environments.
We are looking for an AI/ML Engineer who can build trustworthy datasets, evaluate model and agent performance, and turn measured weaknesses into practical improvements through better data, retrieval and fine-tuning.
You will own the engineering work that helps us answer three questions: Where does the system fail? What should we improve? Does the change perform better on unseen cases?
Strong applied AI/ML skills are essential. Cybersecurity or VAPT experience is a plus. Our security specialists will help define domain labels, review evidence and validate security findings.
What you will work on
- Data pipelines: Transform approved agent traces, tool outputs, expert feedback and verified outcomes into structured, versioned datasets.
- Data quality: Implement cleaning, deduplication, annotation workflows, provenance tracking and sensitive-data handling.
- Evaluation datasets: Create training, validation and held-out test sets, preventing related cases or near-duplicate examples from leaking across splits.
- Model and agent evaluation: Build repeatable evaluations for task completion, tool-use accuracy, evidence quality, false positives, latency and compute cost.
- Failure analysis: Investigate whether failures originate from the model, prompts, retrieval, tools, data or workflow design, and recommend the appropriate improvement.
- Fine-tuning: Run reproducible supervised fine-tuning experiments, including approaches such as LoRA or QLoRA where suitable.
- Retrieval improvement: Evaluate embeddings, rerankers and retrieval pipelines to improve the relevance of the information supplied to agents.
- Experiment tracking: Record dataset versions, configurations, model checkpoints, metrics and failure cases so another engineer can reproduce the results.
- Product integration: Work with the runtime engineer to integrate accepted models and improvements with compatibility tests, regression checks and rollback support.
- Feedback pipelines: Turn reviewed product failures into candidate dataset, evaluation or model improvements before promoting them into a release.What we are looking for
- Strong Python skills and practical experience building data processing or machine learning pipelines.
- Working knowledge of PyTorch and the Hugging Face ecosystem, or comparable tools.
- A hands-on LLM fine-tuning or post-training project covering data preparation, a baseline, held-out evaluation and failure analysis.
- Understanding of training versus inference, overfitting, generalisation, loss functions and evaluation metrics.
- Familiarity with dataset quality, annotation, deduplication and data leakage prevention.
- Understanding of LLM prompting, structured outputs, retrieval and tool use.
- Ability to work with Linux, Git, GPU environments and reproducible experiment configurations.
- Ability to explain why a change improved results—or why it should be rejected.A well-executed research project, substantial personal project or implemented coursework can demonstrate relevant skills. Be ready to explain your contribution and the limitations of your results.
Good to have
- Familiarity with cybersecurity, VAPT workflows, vulnerability reports or security datasets.
- Experience with preference optimisation, DPO, reinforcement learning or distillation.
- Exposure to PEFT, TRL, experiment tracking and dataset versioning tools.
- Experience with model quantisation, GPU memory optimisation or distributed training.
- Experience evaluating RAG systems, embeddings, rerankers or agent trajectories.
- Familiarity with private model deployment and inference serving.You do not need experience with every technology listed.
What you will gain
- Hands-on access to SecNinjaz’s on-premises NVIDIA H200 GPU infrastructure for planned product training, fine-tuning, evaluation and inference workloads.
- An opportunity to deepen your expertise through real experiments and measurable product outcomes.
- Collaboration with cybersecurity specialists who can help turn domain expertise into high-quality datasets and evaluations.
- Ownership of the improvement process, from identifying a failure to evaluating and integrating a better solution.Your initial contribution
Establish a reviewed dataset and evaluation baseline, then run a bounded model-improvement experiment. Compare the candidate against the baseline on unseen cases and make an evidence-based recommendation on adoption.
How to apply
Apply through LinkedIn with your CV and a relevant project link,
- We are especially interested in how you measured improvement, prevented data leakage and investigated failures.
Or
email your CV to nitin@secninjaz.com with a short technical write-up explaining your dataset, fine-tuning approach, evaluation results and personal contribution.and completed and ongoing AI/agentic AI projects, your specific contributions, current CTC, expected CTC, and notice period or earliest date you can join SecNinjaz.