Skip to content

lab

InvoiceOps

No value the LLM produces can reach a policy decision, applied to invoice approval

What it does

A governed accounts-payable intake pipeline: extract → policy-check → confidence-route → human-gate → audit. Drop an invoice PDF in a watched folder and it comes out as auto-record, needs-approval or rejected, with a human-readable reason on every decision and a per-field confidence on every extracted value.

Everyone building invoice automation hits the same question. How much of the decision do you let the model make? For money movement the honest answer is none of it, and this is what taking that seriously looks like when it is enforced architecturally rather than by convention.

Walkthrough on request

Screens

One invoice's audit trail: the pipeline stage by stage, five policy checks each with a plain-English reason, per-field extraction confidences, and a vendor-resolution panel.
The claim in the tagline, on screen. The vendor-resolution panel is labelled the one place an LLM is involved, and underneath it: stored for this trail and never read by the policy engine. PolicyInput has no field it could be assigned to. Every policy check carries its arithmetic, not a verdict. View full size
The intake feed: fourteen invoices as cards, each with a route badge and a confidence band, several showing OCR damage in the vendor names.
Fourteen synthetic invoices, routed. The damage is the point. Orre1l Partners with a digit for an l, a category reading faci1ities, one card noting it was printed as “Acme Logistics Lt”, and two duplicate submissions that resolved to no vendor at all and stopped on two failed rules each. View full size
The exception queue: each held invoice states why it stopped, with approve and reject buttons and an optional note field.
Why this stopped, every time, in a sentence a controller can act on. A PO that was never raised, or an invoice billing £947.64 against one authorising £740.58, over by 28.0% on a 1.0% tolerance. Note the second card: confidence was High at 0.85 and it was gated anyway, because confidence and policy are separate gates. View full size
Confidence analytics: per-field histograms against the auto-record threshold, with four fields marked critical.
Per-field confidence against the 0.85 auto-record threshold, and the reason only four fields are marked critical is written on the page: a hazy line-item description does not escalate an invoice, a hazy amount does. View full size

Stack

  • Python
  • LangGraph
  • FastAPI
  • Azure Bicep

All data in this project is synthetic.