Senior Applied AI Engineer — LLM Reasoning & Evaluation
About the ProjectThe Ops Layer is building HELIOS, an AI-powered operational intelligence platform that transforms organizational questionnaire responses into structured, evidence-grounded decision intelligence.
HELIOS analyzes information from multiple organizational stakeholders to identify recurring operational decisions, decision ownership, authority boundaries, escalation conditions, dependencies, and governance gaps.
An initial implementation already exists, including data ingestion, LLM processing, claim extraction, semantic normalization, evidence lineage, and structured Excel output.
We are looking for an experienced Applied AI Engineer to evaluate, strengthen, and advance the existing system, with particular emphasis on semantic reasoning accuracy, reliability, and production readiness.
Key Responsibilities
- Review and assess the existing AI processing architecture and Python backend.
- Improve extraction and interpretation of source-grounded organizational claims.
- Develop reliable recurring-decision identification and semantic consolidation.
- Distinguish operational decisions from activities, recommendations, inputs, prerequisites, and authority statements.
- Preserve decision ownership, approval boundaries, conditions, numerical thresholds, and exception paths.
- Prevent hallucinated decisions, unsupported authority assignments, and incorrect organizational diagnoses.
- Implement evidence-validation mechanisms and source-level traceability.
- Build automated LLM evaluation, regression testing, and quality assurance frameworks.
- Optimize pipeline performance, reliability, and deployment stability.
- Maintain clear technical documentation and reproducible development environments.
Required Technical Experience
- Strong Python backend engineering, preferably FastAPI and SQLAlchemy.
- Hands-on experience developing and deploying LLM-powered applications.
- Experience with OpenAI, Anthropic, or comparable commercial LLM APIs.
- Structured extraction, JSON schemas, function calling, and output validation.
- Semantic classification, entity resolution, and information normalization.
- Multi-stage AI workflows and reasoning pipelines.
- LLM evaluation frameworks, gold-standard datasets, and regression testing.
- Experience with relational databases, REST APIs, Git, and cloud deployment.
- Ability to investigate complex semantic failures systematically.
Strongly Preferred
- Experience with knowledge graphs or graph-based reasoning.
- Document intelligence, enterprise AI, or decision-support systems.
- Evidence-grounded generation and hallucination mitigation.
- Human-in-the-loop review architectures.
- Model benchmarking, automated evaluation, and deterministic validation.
What Success Looks LikeThe immediate objective is to improve HELIOS's ability to:
- Correctly identify recurring organizational decisions from multiple respondents.
- Consolidate semantically equivalent decisions without merging distinct authority scopes.
- Preserve decision-specific inputs, conditions, thresholds, and dependencies.
- Distinguish formal authority, actual practice, exceptions, and unresolved information.
- Ensure every published interpretation is supported by traceable evidence.
- Produce consistent, independently verifiable results across diverse organizational scenarios.The first milestone will focus on evaluating and improving the existing reasoning engine against predefined acceptance tests.
We prioritize demonstrated semantic correctness over pipeline complexity or feature volume.
Ideal CandidateAn engineer who has built AI systems where accuracy, traceability, and consistent interpretation are critical.
We value strong engineering judgment, transparent communication, reliable delivery practices, and the ability to identify root causes rather than repeatedly patch individual outputs.
Candidates should be comfortable taking technical ownership of an existing codebase and proposing architectural improvements when justified.
Selection ProcessShortlisted candidates will participate in a technical discussion followed by a paid, fixed-scope evaluation involving a representative organizational reasoning scenario.
Please share:
- Relevant AI/LLM projects you have built.
- Your specific technical contributions.
- GitHub repositories or technical work samples, where available.
- Experience building production-grade AI applications.
- Availability and indicative contract rates.
- Apply: vishnu@theopslayer.com