Skip to content

LedgerGuard

An agent that can prove what it did, under seven governance checks, applied to bank reconciliation

  • governance
  • auditability
  • human-in-the-loop
  • Azure AI Foundry

38

decisions in one reconciliation

24 auto-posted, 14 paused for a human. The split is the point: the system declines to decide where it should not.

45

OTel spans read back through the collector

7

governance checks, each independently demoable offline

4

real scorer bugs found by running the eval suite

in short

LedgerGuard reconciles a bank statement and can reconstruct every decision it made. Confidence is arithmetic over five named ledger features, never a model's self-report. Seven governance checks each run offline as one command. Check six lowers one threshold and shows an accuracy gate passing a configuration that silently stops reviewing short payments.

the problem

Why this is hard

Reconciliation is the ideal first job for a finance agent: high volume, rule-shaped, genuinely tedious. It is also the worst possible place for a confident guess. Marking an invoice settled when it was short-paid quietly writes off money somebody still owes. And the mistake is invisible, because the books balance, the queue is short, and everybody is pleased.

So the whole system is built around one claim: every decision it makes is reconstructable, and every decision it declines to make is on the record. A reviewer asking "why 0.62?" gets an itemised answer that adds up, because the confidence score is arithmetic over five named features read from the ledger rather than a number a model reported about itself.

constraints

  • All data is synthetic, generated from seeded templates. No real financial data is used anywhere, including in the security payloads.
  • The same container must run locally and under Azure AI Foundry. Not a near-copy, the same image, with one switch selecting provider, secrets, trace exporter and approval producer.
  • A match cannot reach the database without a reason and citations that resolve to evidence the agent actually looked up.
  • Every check must run offline, with no API key and no cloud account.

where else this pattern applies

The transferable artefact is an audit trail that is an architecture rather than a log, plus a reviewer console where a human decision is captured as data. Clinical decision support has the same requirement: a recommendation is worthless unless a reviewer can see what evidence produced it and record why they overrode it.

Legal document review needs the two-floor pattern specifically. The one-number config change here demonstrates why a single confidence threshold collapses under a shifting document mix. The same failure a contract-triage agent hits when a new counterparty template arrives.

Any regulated workflow where a decision must be defensible months later. Benefits eligibility, KYC review, export-control screening. Needs the governance checks to run as gates in the graph rather than as assertions in a test suite.

architecture

How it runs

A LangGraph graph loads a period, fetches data over MCP, proposes matches, scores confidence arithmetically, then routes: at or above 0.85 it finalises; below that it emits an approval request and interrupts, checkpointing to Postgres. A signed callback resumes it. The Reviewer Console's Approve button does not apply a decision directly. It signs the same contract a Teams card would and posts it over loopback, deliberately the long way round.

  1. 01Load periodplan

    2025-Q1 against ledgerlab's realistic profile, seed 42. 38 transactions to decide.

  2. 02Fetch over MCPtool

    Seven MCP tools over Streamable HTTP. A committed fixture export sits behind the same interface as a fallback, a fallback, not a mock.

  3. 03Score confidencepolicy

    Arithmetic over five named features read from the ledger. There is no code path from prose to the number, which is why the injection corpus scores 0% attack success structurally rather than by filtering.

  4. 04Route on thresholdroute

    ≥0.85 auto-posts. 0.60 to 0.85 and below both go to a human. 24 auto-posted; 14 paused.

  5. 05Abstain where evidence runs outhalt

    Three of the fourteen propose no target at all. An escalation that names nothing beats a guess that names something.

  6. 06Interrupt and checkpointgate

    The graph interrupts and checkpoints to Postgres. Duplicate delivery of the same decision returns replayed: true and resumes nothing.

  7. 07Resume and completecommit

    Fourteen decisions posted through the console's own signed callback; the last one resumed the graph and the run completed.

  8. 08Audit exportaudit

    38 rows, every one with its trace id, corrected rows carrying the agent's original proposal alongside the reviewer's choice.

  • passed
  • awaiting a person
  • stopped

decisions

What was considered, and why it was ruled out

Every one of these had a reasonable alternative. The alternative is named.

ADR-01

Confidence is arithmetic over ledger features, never a model's self-report

Alternative
Ask the model for a confidence score, which is what most agent stacks do.
Why not
A reviewer asking why 0.62 needs an answer that adds up, and a self-reported score cannot give one. It also turned out to be the security property: the prompt-injection corpus scores 0% attack success and 0% false-block rate not because a filter catches attacks, but because a memo is not one of the five features, so there is no code path from prose to the number.

ADR-02

The Approve button posts the Teams contract to itself over loopback

Alternative
Have the console apply the decision directly, one fewer hop, obviously simpler.
Why not
The spec calls the Logic-Apps-to-LangGraph resume the highest-risk integration in the project. Routing the console's own button through the identical signed contract means every click in local development exercises the Teams path, so that risk got spent in week one instead of on a deploy day. A contract test asserts the Logic App's outgoing body against the same Pydantic model, which is how a 422-on-every-card bug was caught before deployment rather than after.

ADR-03

Gate the deploy on two floors, not one

Alternative
Gate on matching accuracy, the metric everybody reports.
Why not
Check six demonstrates why. A degraded config differing by one number. The auto threshold, 0.85 to 0.75, to shorten a queue reviewers had complained about. Holds MatchF1 at a perfect 1.0000 while EscalationScore falls to 0.6700. Nothing errors, every decision still carries grounded citations, and the queue halves. What actually happened is that all seven short-payment escalations stopped being reviewed, so short-paid invoices are now marked settled with no human involved. A gate on matching accuracy alone would have passed it.

ADR-04

Enforce dependency direction in the image build, not by convention

Alternative
A lint rule or a code-review norm about which package imports which.
Why not
The agent image cannot import the companion or the benchmark, and the companion cannot import LangGraph. Because the builds do not contain them. A convention degrades the first time someone is in a hurry; a build failure does not.

evaluation

How it was measured

Running the eval suite found four real defects in the confidence scorer that no unit test in the repository would have caught. The first run scored MatchF1 0.4827 and EscalationScore 0.4365. Each defect now has a named regression test. That is why the gate earns its place rather than decorating the pipeline.

metricresultnote
MatchF11.0000100 cases × k=4, mode: offline
EscalationScore1.0000floor is 0.90
Traces38/3845 spans read back through the collector
Prompt-injection run0% attack success0% false-block, 0 suppressed reviews, 120 cases
Tests367agent 171 · backend 171 · frontend 25
Azure IaC8 templates, 0 errorspinned standalone Bicep 0.31.92

stack

Built with

  • Python
  • LangGraph
  • FastAPI
  • Postgres
  • OpenTelemetry
  • Azure Bicep
  • React