Walkthroughs and screens captured from the real applications. Everything here was rendered by the project it belongs to. No mockups, no concept art. Projects that have not been captured yet are simply absent rather than represented by something that looks like a screenshot.
13 of 16 projects have a walkthrough or screens · all data synthetic
Tool-grounded answers over a deterministic subscription revenue engine
The agent refuses the first question because no data is loaded, then the sample goes in and the same question is answered from the computed schedule rather than from the model.
Vision extraction checked by arithmetic, with bounded repair and a second check on the repair itself
An invoice where the extraction is plausible and wrong. The arithmetic catches three violations, the repair loop runs, and a second check watches the repair for gaming the sums.
Governed RAG where weak retrieval refuses instead of guessing, applied to a finance policy manual
The same question asked as three roles. Guest and staff are refused with the restricted passages never retrieved, and the controller is answered with every claim carrying a citation.
1 / 3 · screens
One question, three roles, side by side. Guest and staff are refused with six passages withheld by the label filter; the controller is answered from two restricted sources. The line above the panels is the design decision. The filter runs inside the retriever, so restricted passages are never read rather than read and then hidden. View full size
Refusal is a node, not an apology appended to an answer. It scores the retrieval (0.46 against a 0.80 threshold), says plainly that it will not infer, lists the closest sections under a heading insisting they are not an answer, and logs the question as a possible gap in the manual. View full size
The same machinery when retrieval is sufficient (0.94). Every claim carries a numbered citation and the panel states the rule: uncited claims are removed before render. The prose is blunt because this ran with no model credential. Composition falls back to an extractive stub, while citation, label filtering and refusal stay deterministic and unaffected. View full size
Production agent failures turned into versioned eval cases, applied to finance agent traces
Twelve fixture files in two trace formats dropped in live. Malformed records are reported rather than dropped and PII is redacted before it reaches a screen, then runs are labelled in three keystrokes and exported as plain JSONL and a pytest suite.
1 / 3 · screens
The labelling screen. Verdict on R/W/P, failure tags on keys 1 through 7, commit on enter. And a median-seconds-per-label readout with a target of under ten, because a labelling tool that is slow does not get used. View full size
Labelled runs become versioned cases with assertions pre-filled from what the trace actually did. View full size
Framework-neutral output you own: cases.jsonl, a generated pytest suite, an optional Promptfoo config. Nothing is locked to this tool. View full size
A prompt-injection corpus with benign twins, so it measures discrimination not blocking, applied to finance documents
The same sixty documents through a naive agent and a guarded one. Attack success falls from 36.7% to 1.7% with no increase in false blocks, and the run exports as a timestamped report tied to a corpus hash.
1 / 2 · screens
The compliance report. Attack success and false-block rate side by side, because a corpus that only measures blocking tells you half the story. View full size
Per-vector results across document types, with the benign twins scored alongside. View full size
Analysis with receipts, where the model never sees the source data, applied to financial statements
Statements computed from the profile, then written up live by the model without it ever seeing the statements. Every figure in the commentary is a chip that opens the computation behind it.
Confidence, gate, learn. The simplest honest agent loop, applied to expense categorisation
A month of card transactions categorised, with anything under the confidence gate held for review. An override is written to vendor memory, so the next run covers more from memory and calls the model less.
A monthly finance pack narrated by a model that cannot invent a number, applied to management reporting
Assemble a period, review the draft beside the figures each section was allowed to cite, watch it refuse to issue while gaps are open, waive them with a reason, sign, then verify the archived pack against its hash.
No value the LLM produces can reach a policy decision, applied to invoice approval
1 / 4 · screens
The claim in the tagline, on screen. The vendor-resolution panel is labelled the one place an LLM is involved, and underneath it: stored for this trail and never read by the policy engine. PolicyInput has no field it could be assigned to. Every policy check carries its arithmetic, not a verdict. View full size
Fourteen synthetic invoices, routed. The damage is the point. Orre1l Partners with a digit for an l, a category reading faci1ities, one card noting it was printed as “Acme Logistics Lt”, and two duplicate submissions that resolved to no vendor at all and stopped on two failed rules each. View full size
Why this stopped, every time, in a sentence a controller can act on. A PO that was never raised, or an invoice billing £947.64 against one authorising £740.58, over by 28.0% on a 1.0% tolerance. Note the second card: confidence was High at 0.85 and it was gated anyway, because confidence and policy are separate gates. View full size
Per-field confidence against the 0.85 auto-record threshold, and the reason only four fields are marked critical is written on the page: a hazy line-item description does not escalate an invoice, a hazy amount does. View full size