Skip to content

LedgerLens

An exception workbench where an uncited hypothesis is deleted in code, applied to the breaks AI matching cannot close

  • agents
  • grounding
  • human-in-the-loop
  • evals
A narrated run through the exception workbench: import, the queue, opening a transaction, and reading the evidence behind a match before confirming it.

96%

top-1 root-cause accuracy

on the live model path, over synthetic reconciliation exceptions. Higher is better.

100%

citations resolving to in-context records (58/58)

0

uncited hypotheses that reached the database, ever

$0.00098

cost per exception investigated

in short

Residual reconciliation exceptions fail for lack of context, not algorithm quality, so a human investigates each one cold. LedgerLens has a LangGraph agent gather the related invoices, prior payments and notes, then draft root-cause hypotheses that must cite evidence or be deleted before storage. The reviewer confirms in one keystroke.

the problem

Why this is hard

AI matching engines stall somewhere between 85% and 92%. The exceptions they leave behind do not fail because the algorithm is weak. They fail because the deciding context is somewhere else: a prior payment, an invoice under a different name, a bookkeeper's note about a dispute. So a human investigates each one from scratch.

LedgerLens splits that differently. The AI does the investigation; the human does the judgment. A LangGraph agent gathers the context pack, drafts one to three root-cause hypotheses, and each one must cite evidence that resolves to a record it actually saw. A reviewer confirms with one keystroke instead of digging.

constraints

  • Matching itself is an explicit non-goal. This tool starts where the matching engine gives up.
  • A per-batch cost cap is checked before each call, so it is a ceiling rather than a target.
  • Ambiguous dates are flagged, never guessed, 03/04/2025 is not silently resolved.
  • unknown is a first-class answer, styled neutrally and measured, and the agent is rewarded for it when context is genuinely absent.

where else this pattern applies

Every high-volume classifier leaves a residue it cannot close, and the interesting product is almost always the queue for that residue rather than the model. Content moderation appeals, fraud-alert triage and medical-coding review are the same shape: a confident bulk path plus a human workbench for what it refused.

The enforced rule. An uncited hypothesis is deleted before render, in code rather than by prompt. Transfers to any assistive tool whose suggestions a professional is accountable for. A radiology second-read or a legal citation check fails in exactly the way an uncited reconciliation hypothesis does.

The pattern library, where a resolved exception becomes a reusable rule, is how any triage queue stops growing linearly with volume.

architecture

How it runs

A five-node LangGraph graph: load the exception, gather context, hypothesise, ground-check, rank. With one re-ask if grounding drops everything. The ground check is a deterministic node, not a prompt instruction, and the runner is the only code that writes a hypothesis row, writing only from the graph's final state. That is what makes "uncited hypotheses never reach the UI" structural rather than conventional.

  1. 01Load exceptionplan

    One residual transaction the matching engine could not settle.

  2. 02Gather context packtool

    Related invoices, prior payments from the same counterparty, and the notes that decide disputes.

  3. 03Draft 1 to 3 hypothesesroute

    Each is {statement, root_cause_type, confidence, evidence[]} across seven root causes.

  4. 04Ground check, deterministicpolicy

    Three checks: it must cite something, every cited id must resolve to a real stored record, and that record must have been in this exception's context pack.

  5. 05Uncited drafts are deletedhalt

    Not flagged, not down-ranked. Deleted. A hypothesis citing a plausible-looking id it never saw is a lucky guess, and it goes too.

  6. 06Rank and storecommit

    The runner writes only from the graph's final state, so the database cannot hold an ungrounded claim.

  7. 07Reviewer verdict, immutableaudit

    C confirm, X correct, R reject. A second verdict on the same exception is a 409, not an overwrite.

  • passed
  • awaiting a person
  • stopped

decisions

What was considered, and why it was ruled out

Every one of these had a reasonable alternative. The alternative is named.

ADR-01

Ground-check in a deterministic node, not in the prompt

Alternative
Instruct the model to cite its sources and validate the citations afterwards, flagging bad ones.
Why not
Prompt instructions are requests. The check is three lines of Python that delete an uncited draft before it is stored, rendered or counted. A test proves it at the database level: a deliberately sabotaging investigator emits a 0.99-confidence claim with no evidence and a 0.97-confidence claim citing an invented invoice, and both tables come back empty.

ADR-02

Make unknown a rewarded answer rather than a failure state

Alternative
Always produce a best-guess root cause and let confidence carry the uncertainty.
Why not
Some payments genuinely have no explanation in the books. An agent that always answers scores well by confidently matching everything to something, which is exactly the behaviour that makes a reconciliation tool dangerous. Measuring unknown honesty and over-abstention as separate metrics means refusing correctly and refusing lazily are distinguishable.

ADR-03

Let heavily truncated counterparty names form their own pattern group

Alternative
Normalise harder so ACME 4471 and ACMESUPPLIES land in the same group as a human would expect.
Why not
No string rule rejoins those without also merging Acme Supplies with Acme Logistics, and one misleading pattern is worse than two split ones. Retrieval works around it with a prefix fallback, because there a wrong candidate is only a distractor the agent rejects. Pattern grouping deliberately does not, and a test named after the limitation asserts the behaviour.

ADR-04

Forbid the deterministic fallback from reading the answer key, by test

Alternative
Trust that the fake LLM used for offline runs is written honestly.
Why not
The heuristic path is what CI runs and what every committed baseline number comes from. If it could see ground truth, every one of those numbers would be theatre. A test asserts it cannot, so the baseline is honest rather than rigged. Which is what makes the live model's improvement over it meaningful.

evaluation

How it was measured

The live model path has been run end to end. The only flagship here where that is true. Grounding held against a real model, which is the first adversary capable of inventing a plausible invoice id. The eval suite also caught two real bugs during the build, and both fixes were to the data rather than the classifier.

metricresultnote
Top-1 root cause (live model, 25-case demo set)96.0%24/25. Heuristic baseline was 84.0%
Top-1 root cause (heuristic, 30-case suite)93.3%different data and difficulty mix, not comparable to the above
Evidence validity100.0%structural; gates CI unconditionally
unknown honesty100.0%refused to guess on all 7 unanswerable cases
Over-abstention0.0%never abandoned an answerable case
Tests179no network, no spend

screens

What it looks like running

Captures of the real thing. Projects without a capture show none, nothing here is a mockup.

The review queue: 25 exceptions with root-cause chips, confidence pills and a batch progress bar.
The review queue. Root-cause chips, confidence pills, and a batch run in progress over the WebSocket. View full size
An exception detail view in three panes: the transaction, ranked hypotheses with evidence chips, and the evidence panel with matched fields highlighted.
One exception. Ranked hypotheses, each with the evidence chips it cited. Click one and the source record opens with the matching fields highlighted. View full size
The pattern library, grouping confirmed resolutions by root cause and counterparty.
The pattern library. Confirmed resolutions group by root cause and counterparty, so the third occurrence arrives with 'seen 3 times before' attached. View full size
The resolution report, showing the AI top-one agreement tile above an immutable verdict log.
The resolution report. Note the caveat in the limits below: the agreement figure only means something once corrections are in the log. View full size

stack

Built with

  • Python
  • LangGraph
  • FastAPI
  • SQLite
  • React
  • Vite
  • WebSockets