Senior Applied AI Engineer
Global Growth – Global BD Agent Suite:
Build the intelligence and behaviour of the Global BD agents on Azure. That means retrieval, prompting, tool use, grounded and cited generation, and structured outputs for document-heavy pitch workflows: multi-hundred-page RFP packs, Excel DTT grids, POAPs, proofing of Word and PowerPoint files, and case studies. Every agent is built on the shared backbone so its components can be reused by the rest of the suite.
Mandatory focus:
Azure AI Search retrieval tuning (vector, hybrid, semantic ranker, chunking); Azure OpenAI; tool/function calling with structured outputs in Foundry Agent Service or Semantic Kernel; source-cited generation with confidence and gap flags; RAG and agent evaluation.
Key Responsibilities:
- Build agent logic for the MVP pitching cluster and later agents: RFP interrogation, first-draft RFP responses, POAP starter, proofing suggestions and DTT grid pre-population.
- Design retrieval on Azure AI Search: chunking strategy, embeddings, hybrid and semantic ranking, metadata filters and ACL-aware filtering, tuned against a measured evaluation set.
- Implement grounded generation in which every output separates retrieved fact from generated narrative, cites sources and flags missing, stale or contradictory information instead of filling gaps.
- Implement tool / function calling and structured outputs (JSON schemas) so agents can read templates, fill grids, call live CRM or document tools and hand work to the next agent.
- Handle document-heavy I/O by parsing multi-document PDF, Word, Excel and PowerPoint inputs and producing editable Word, Excel and PowerPoint outputs and tracked-change style suggestions.
- Build reusable components such as template interpreter, glossary and style checker, citation formatter and prompt templates that other agents can share.
- Own evaluation by building golden datasets with BD SMEs and measuring groundedness, relevance, citation accuracy and task completion; iterate prompts and retrieval against the numbers.
- Optimise latency, token usage and cost per workflow run; separate heavy file ingestion from quick Q&A paths.
Required Skills & Expected Capability:
- GenAI / RAG: Production RAG with embeddings, chunking, hybrid retrieval, re-ranking and citation; has debugged poor answers by fixing retrieval rather than only rewriting prompts.
- Agents / Tool Use: Function / tool calling, structured outputs, multi-step agents, agent handoffs and human-in-the-loop steps using Foundry Agent Service, Semantic Kernel or an equivalent framework.
- Azure AI: Azure OpenAI, Azure AI Search, Azure AI Foundry (agents, prompt flow, evaluations); Azure AI Document Intelligence is a plus.
- Document Processing: Parsing and generating Office and PDF content (python-docx, openpyxl, python-pptx, PDF parsers); table and form extraction from long documents.
- Evaluation: Building eval sets and automated metrics (groundedness, relevance, faithfulness) with tools such as Foundry evaluations, Ragas, promptfoo or DeepEval.
- Engineering: Strong Python, Git, API development (FastAPI), unit testing, CI/CD basics.
You will fit right in if you have:
- 5–8 years in software or ML engineering, with 2+ years building LLM / RAG applications that real users rely on.
- Has shipped at least one tool-using agent or multi-step LLM workflow, not only a chatbot.
- Can explain, with numbers, how they improved retrieval quality on a real project.
- Works directly with non-technical SMEs to turn their review comments into eval cases.