LedgerGuard
An agent that can prove what it did, under seven governance checks, applied to bank reconciliation
- governance
- auditability
- human-in-the-loop
- Azure AI Foundry
38
decisions in one reconciliation
24 auto-posted, 14 paused for a human. The split is the point: the system declines to decide where it should not.
45
OTel spans read back through the collector
7
governance checks, each independently demoable offline
4
real scorer bugs found by running the eval suite
in short
LedgerGuard reconciles a bank statement and can reconstruct every decision it made. Confidence is arithmetic over five named ledger features, never a model's self-report. Seven governance checks each run offline as one command. Check six lowers one threshold and shows an accuracy gate passing a configuration that silently stops reviewing short payments.
the problem
Why this is hard
Reconciliation is the ideal first job for a finance agent: high volume, rule-shaped, genuinely tedious. It is also the worst possible place for a confident guess. Marking an invoice settled when it was short-paid quietly writes off money somebody still owes. And the mistake is invisible, because the books balance, the queue is short, and everybody is pleased.
So the whole system is built around one claim: every decision it makes is reconstructable, and every decision it declines to make is on the record. A reviewer asking "why 0.62?" gets an itemised answer that adds up, because the confidence score is arithmetic over five named features read from the ledger rather than a number a model reported about itself.
constraints
- All data is synthetic, generated from seeded templates. No real financial data is used anywhere, including in the security payloads.
- The same container must run locally and under Azure AI Foundry. Not a near-copy, the same image, with one switch selecting provider, secrets, trace exporter and approval producer.
- A match cannot reach the database without a reason and citations that resolve to evidence the agent actually looked up.
- Every check must run offline, with no API key and no cloud account.
where else this pattern applies
The transferable artefact is an audit trail that is an architecture rather than a log, plus a reviewer console where a human decision is captured as data. Clinical decision support has the same requirement: a recommendation is worthless unless a reviewer can see what evidence produced it and record why they overrode it.
Legal document review needs the two-floor pattern specifically. The one-number config change here demonstrates why a single confidence threshold collapses under a shifting document mix. The same failure a contract-triage agent hits when a new counterparty template arrives.
Any regulated workflow where a decision must be defensible months later. Benefits eligibility, KYC review, export-control screening. Needs the governance checks to run as gates in the graph rather than as assertions in a test suite.
architecture
How it runs
A LangGraph graph loads a period, fetches data over MCP, proposes matches, scores confidence arithmetically, then routes: at or above 0.85 it finalises; below that it emits an approval request and interrupts, checkpointing to Postgres. A signed callback resumes it. The Reviewer Console's Approve button does not apply a decision directly. It signs the same contract a Teams card would and posts it over loopback, deliberately the long way round.
01Load periodplan
2025-Q1 against ledgerlab's realistic profile, seed 42. 38 transactions to decide.
02Fetch over MCPtool
Seven MCP tools over Streamable HTTP. A committed fixture export sits behind the same interface as a fallback, a fallback, not a mock.
03Score confidencepolicy
Arithmetic over five named features read from the ledger. There is no code path from prose to the number, which is why the injection corpus scores 0% attack success structurally rather than by filtering.
04Route on thresholdroute
≥0.85 auto-posts. 0.60 to 0.85 and below both go to a human. 24 auto-posted; 14 paused.
05Abstain where evidence runs outhalt
Three of the fourteen propose no target at all. An escalation that names nothing beats a guess that names something.
06Interrupt and checkpointgate
The graph interrupts and checkpoints to Postgres. Duplicate delivery of the same decision returns replayed: true and resumes nothing.
07Resume and completecommit
Fourteen decisions posted through the console's own signed callback; the last one resumed the graph and the run completed.
08Audit exportaudit
38 rows, every one with its trace id, corrected rows carrying the agent's original proposal alongside the reviewer's choice.
- passed
- awaiting a person
- stopped
decisions
What was considered, and why it was ruled out
Every one of these had a reasonable alternative. The alternative is named.
ADR-01
Confidence is arithmetic over ledger features, never a model's self-report
- Alternative
- Ask the model for a confidence score, which is what most agent stacks do.
- Why not
- A reviewer asking why 0.62 needs an answer that adds up, and a self-reported score cannot give one. It also turned out to be the security property: the prompt-injection corpus scores 0% attack success and 0% false-block rate not because a filter catches attacks, but because a memo is not one of the five features, so there is no code path from prose to the number.
ADR-02
The Approve button posts the Teams contract to itself over loopback
- Alternative
- Have the console apply the decision directly, one fewer hop, obviously simpler.
- Why not
- The spec calls the Logic-Apps-to-LangGraph resume the highest-risk integration in the project. Routing the console's own button through the identical signed contract means every click in local development exercises the Teams path, so that risk got spent in week one instead of on a deploy day. A contract test asserts the Logic App's outgoing body against the same Pydantic model, which is how a 422-on-every-card bug was caught before deployment rather than after.
ADR-03
Gate the deploy on two floors, not one
- Alternative
- Gate on matching accuracy, the metric everybody reports.
- Why not
- Check six demonstrates why. A degraded config differing by one number. The auto threshold, 0.85 to 0.75, to shorten a queue reviewers had complained about. Holds MatchF1 at a perfect 1.0000 while EscalationScore falls to 0.6700. Nothing errors, every decision still carries grounded citations, and the queue halves. What actually happened is that all seven short-payment escalations stopped being reviewed, so short-paid invoices are now marked settled with no human involved. A gate on matching accuracy alone would have passed it.
ADR-04
Enforce dependency direction in the image build, not by convention
- Alternative
- A lint rule or a code-review norm about which package imports which.
- Why not
- The agent image cannot import the companion or the benchmark, and the companion cannot import LangGraph. Because the builds do not contain them. A convention degrades the first time someone is in a hurry; a build failure does not.
evaluation
How it was measured
Running the eval suite found four real defects in the confidence scorer that no unit test in the repository would have caught. The first run scored MatchF1 0.4827 and EscalationScore 0.4365. Each defect now has a named regression test. That is why the gate earns its place rather than decorating the pipeline.
| metric | result | note |
|---|---|---|
| MatchF1 | 1.0000 | 100 cases × k=4, mode: offline |
| EscalationScore | 1.0000 | floor is 0.90 |
| Traces | 38/38 | 45 spans read back through the collector |
| Prompt-injection run | 0% attack success | 0% false-block, 0 suppressed reviews, 120 cases |
| Tests | 367 | agent 171 · backend 171 · frontend 25 |
| Azure IaC | 8 templates, 0 errors | pinned standalone Bicep 0.31.92 |
stack
Built with
- Python
- LangGraph
- FastAPI
- Postgres
- OpenTelemetry
- Azure Bicep
- React
Aneeq Khatri