Skip to content

lab

InvoiceAudit

Vision extraction checked by arithmetic, with bounded repair and a second check on the repair itself

An invoice where the extraction is plausible and wrong. The arithmetic catches three violations, the repair loop runs, and a second check watches the repair for gaming the sums.

What it does

InvoiceAudit treats a model extraction as an untrusted proposal. Deterministic arithmetic checks the record, quantified failures guide a bounded repair loop, and the router sends uncertain documents for human review. The web interface shows the source document beside violations and the repair trace.

A separate detector compares the original and repaired records. It flags aggregate-only changes that can satisfy arithmetic without evidence of a corrected source reading. The loop is plain Python, with iteration and token limits plus a no-progress stop; there is no second model acting as the arithmetic judge.

The checked-in evaluation report covers 32 synthetic documents using gpt-5.4-nano-2026-03-17. Document accuracy rises from 78.1% to 81.2%, but auto-accept precision falls from 83.3% to 82.3%. The report marks the precision criterion as failed. Better document accuracy is not evidence that unattended acceptance became safer.

The repair that tried to cheat

Scroll through one invoice. Getting the sums to pass is easy; the point is noticing when a repair passes them without a reason.

Illustrative specimen. Every figure is fictional; the steps follow InvoiceAudit's pipeline. On the source the tax reads 246.00: the misread was there, and the repair changed the total instead.

Stack

  • Python
  • Pydantic
  • OpenAI
  • FastAPI
  • React
  • pytest

All data in this project is synthetic.