
Governed RAG where weak retrieval refuses instead of guessing, applied to a finance policy manual
- Python
- LangGraph
- FastAPI
engineering portfolio
Governed multi-agent systems, proven where a wrong number costs the most. Every project gives the problem, the architecture, the decisions and their alternatives, and the evaluation method with its numbers.
All project data is synthetic. Individual projects document their datasets. Available walkthroughs are collected on the demos page.

Governed RAG where weak retrieval refuses instead of guessing, applied to a finance policy manual

Production agent failures turned into versioned eval cases, applied to finance agent traces

A prompt-injection corpus with benign twins, so it measures discrimination not blocking, applied to finance documents

No value the LLM produces can reach a policy decision, applied to invoice approval

Analysis with receipts, where the model never sees the source data, applied to financial statements

Confidence, gate, learn. The simplest honest agent loop, applied to expense categorisation


Vision extraction checked by arithmetic, with bounded repair and a second check on the repair itself

A monthly finance pack narrated by a model that cannot invent a number, applied to management reporting
These run, and are written up in full. They join the grid above once their screens have been recorded.
The full set of repositories is on GitHub.