LedgerLens
An exception workbench where an uncited hypothesis is deleted in code, applied to the breaks AI matching cannot close
- agents
- grounding
- human-in-the-loop
- evals
96%
top-1 root-cause accuracy
on the live model path, over synthetic reconciliation exceptions. Higher is better.
100%
citations resolving to in-context records (58/58)
0
uncited hypotheses that reached the database, ever
$0.00098
cost per exception investigated
in short
Residual reconciliation exceptions fail for lack of context, not algorithm quality, so a human investigates each one cold. LedgerLens has a LangGraph agent gather the related invoices, prior payments and notes, then draft root-cause hypotheses that must cite evidence or be deleted before storage. The reviewer confirms in one keystroke.
the problem
Why this is hard
AI matching engines stall somewhere between 85% and 92%. The exceptions they leave behind do not fail because the algorithm is weak. They fail because the deciding context is somewhere else: a prior payment, an invoice under a different name, a bookkeeper's note about a dispute. So a human investigates each one from scratch.
LedgerLens splits that differently. The AI does the investigation; the human does the judgment. A LangGraph agent gathers the context pack, drafts one to three root-cause hypotheses, and each one must cite evidence that resolves to a record it actually saw. A reviewer confirms with one keystroke instead of digging.
constraints
- Matching itself is an explicit non-goal. This tool starts where the matching engine gives up.
- A per-batch cost cap is checked before each call, so it is a ceiling rather than a target.
- Ambiguous dates are flagged, never guessed, 03/04/2025 is not silently resolved.
- unknown is a first-class answer, styled neutrally and measured, and the agent is rewarded for it when context is genuinely absent.
where else this pattern applies
Every high-volume classifier leaves a residue it cannot close, and the interesting product is almost always the queue for that residue rather than the model. Content moderation appeals, fraud-alert triage and medical-coding review are the same shape: a confident bulk path plus a human workbench for what it refused.
The enforced rule. An uncited hypothesis is deleted before render, in code rather than by prompt. Transfers to any assistive tool whose suggestions a professional is accountable for. A radiology second-read or a legal citation check fails in exactly the way an uncited reconciliation hypothesis does.
The pattern library, where a resolved exception becomes a reusable rule, is how any triage queue stops growing linearly with volume.
architecture
How it runs
A five-node LangGraph graph: load the exception, gather context, hypothesise, ground-check, rank. With one re-ask if grounding drops everything. The ground check is a deterministic node, not a prompt instruction, and the runner is the only code that writes a hypothesis row, writing only from the graph's final state. That is what makes "uncited hypotheses never reach the UI" structural rather than conventional.
01Load exceptionplan
One residual transaction the matching engine could not settle.
02Gather context packtool
Related invoices, prior payments from the same counterparty, and the notes that decide disputes.
03Draft 1 to 3 hypothesesroute
Each is {statement, root_cause_type, confidence, evidence[]} across seven root causes.
04Ground check, deterministicpolicy
Three checks: it must cite something, every cited id must resolve to a real stored record, and that record must have been in this exception's context pack.
05Uncited drafts are deletedhalt
Not flagged, not down-ranked. Deleted. A hypothesis citing a plausible-looking id it never saw is a lucky guess, and it goes too.
06Rank and storecommit
The runner writes only from the graph's final state, so the database cannot hold an ungrounded claim.
07Reviewer verdict, immutableaudit
C confirm, X correct, R reject. A second verdict on the same exception is a 409, not an overwrite.
- passed
- awaiting a person
- stopped
decisions
What was considered, and why it was ruled out
Every one of these had a reasonable alternative. The alternative is named.
ADR-01
Ground-check in a deterministic node, not in the prompt
- Alternative
- Instruct the model to cite its sources and validate the citations afterwards, flagging bad ones.
- Why not
- Prompt instructions are requests. The check is three lines of Python that delete an uncited draft before it is stored, rendered or counted. A test proves it at the database level: a deliberately sabotaging investigator emits a 0.99-confidence claim with no evidence and a 0.97-confidence claim citing an invented invoice, and both tables come back empty.
ADR-02
Make unknown a rewarded answer rather than a failure state
- Alternative
- Always produce a best-guess root cause and let confidence carry the uncertainty.
- Why not
- Some payments genuinely have no explanation in the books. An agent that always answers scores well by confidently matching everything to something, which is exactly the behaviour that makes a reconciliation tool dangerous. Measuring unknown honesty and over-abstention as separate metrics means refusing correctly and refusing lazily are distinguishable.
ADR-03
Let heavily truncated counterparty names form their own pattern group
- Alternative
- Normalise harder so ACME 4471 and ACMESUPPLIES land in the same group as a human would expect.
- Why not
- No string rule rejoins those without also merging Acme Supplies with Acme Logistics, and one misleading pattern is worse than two split ones. Retrieval works around it with a prefix fallback, because there a wrong candidate is only a distractor the agent rejects. Pattern grouping deliberately does not, and a test named after the limitation asserts the behaviour.
ADR-04
Forbid the deterministic fallback from reading the answer key, by test
- Alternative
- Trust that the fake LLM used for offline runs is written honestly.
- Why not
- The heuristic path is what CI runs and what every committed baseline number comes from. If it could see ground truth, every one of those numbers would be theatre. A test asserts it cannot, so the baseline is honest rather than rigged. Which is what makes the live model's improvement over it meaningful.
evaluation
How it was measured
The live model path has been run end to end. The only flagship here where that is true. Grounding held against a real model, which is the first adversary capable of inventing a plausible invoice id. The eval suite also caught two real bugs during the build, and both fixes were to the data rather than the classifier.
| metric | result | note |
|---|---|---|
| Top-1 root cause (live model, 25-case demo set) | 96.0% | 24/25. Heuristic baseline was 84.0% |
| Top-1 root cause (heuristic, 30-case suite) | 93.3% | different data and difficulty mix, not comparable to the above |
| Evidence validity | 100.0% | structural; gates CI unconditionally |
| unknown honesty | 100.0% | refused to guess on all 7 unanswerable cases |
| Over-abstention | 0.0% | never abandoned an answerable case |
| Tests | 179 | no network, no spend |
screens
What it looks like running
Captures of the real thing. Projects without a capture show none, nothing here is a mockup.




stack
Built with
- Python
- LangGraph
- FastAPI
- SQLite
- React
- Vite
- WebSockets
Aneeq Khatri