services
What I build, and what I teach
Two kinds of work: building governed AI systems for teams who need the output to stand up to review, and teaching the engineering practice behind them. Each one below links to a system on this site that demonstrates it, so the claim is checkable before any conversation starts.
AI agent engineering
Multi-agent systems for workflows where being confidently wrong is expensive. Human gates at the risk boundaries, decisions that can be reconstructed months later, and a refusal when the evidence runs out rather than a plausible guess.
- Architecture and build, from the graph and its tools to the approval path and the audit trail.
- Governance designed in rather than logged after: the component that orchestrates the work cannot also approve it.
- Proven against finance data because that is where the constraints are strictest. Each case study sets out where the same pattern transfers: claims adjudication, procurement approval, clinical decision support, moderation and fraud triage, legal document review.
evidenceCloseOpsLedgerGuardRevLedger
Evaluation and reliability
An agent that worked once is not an agent that works. Benchmarks that run every case repeatedly and publish the spread, and regression suites built from the failures a system has already had.
- Eval harnesses that hold the model constant and test the configuration around it.
- Turning production traces into versioned cases, so a failure that has happened once cannot quietly return.
evidenceFinAgent-EvalsTrace2Evals
Finance automation advisory
Month-end close, bank reconciliation, invoice approval, expense categorisation. The accounting judgment rather than the engineering: what a controller will sign off, and what no amount of model quality will make acceptable.
- Where automation is safe, where it is not, and what has to be true before a number posts without a person.
- Reviewing an existing pipeline against the governance a controller would actually ask for.
- Scoped by someone who reconciled the books before automating them, so the plan accounts for what the domain will not tolerate.
Corporate training and workshops
Hands-on sessions for engineering teams shipping AI features: agent architecture, tool calling, evaluation, and the governance that decides whether a feature survives contact with production.
- Built the way the courses are built, around building and breaking a real system rather than slides.
New as a corporate offering. The teaching record behind it is national programmes rather than companies: around a hundred instructors led at the Governor Sindh Initiative, and courses taught at PIAIC and Saylani.
evidenceTeaching
AI engineering training for individuals
The material taught at PIAIC, Saylani and the Governor Sindh Initiative, for engineers moving into agentic AI. Build, break, test and ship a real workflow, rather than assembling a portfolio of demos.
- Multi-agent architectures, tool calling, retrieval, evaluation, and the reliability practice that separates a demo from a system.
evidenceTeaching
working together
Where this stands today
Stated rather than implied, because availability is the first thing a reader is trying to work out.
I am open to fully remote roles worldwide, and that remains the primary route. The work above is available alongside it.
Every project linked from this page uses synthetic data. The systems are real and were run to produce the recordings and screenshots on this site, but no client data appears anywhere in them.
Aneeq Khatri