Skip to content

demos

The systems, running

Walkthroughs and screens captured from the real applications. Everything here was rendered by the project it belongs to. No mockups, no concept art. Projects that have not been captured yet are simply absent rather than represented by something that looks like a screenshot.

13 of 16 projects have a walkthrough or screens · all data synthetic

Tool-grounded answers over a deterministic subscription revenue engine

The agent refuses the first question because no data is loaded, then the sample goes in and the same question is answered from the computed schedule rather than from the model.

Vision extraction checked by arithmetic, with bounded repair and a second check on the repair itself

An invoice where the extraction is plausible and wrong. The arithmetic catches three violations, the repair loop runs, and a second check watches the repair for gaming the sums.

Governed RAG where weak retrieval refuses instead of guessing, applied to a finance policy manual

The same question asked as three roles. Guest and staff are refused with the restricted passages never retrieved, and the controller is answered with every claim carrying a citation.

Production agent failures turned into versioned eval cases, applied to finance agent traces

Twelve fixture files in two trace formats dropped in live. Malformed records are reported rather than dropped and PII is redacted before it reaches a screen, then runs are labelled in three keystrokes and exported as plain JSONL and a pytest suite.

A prompt-injection corpus with benign twins, so it measures discrimination not blocking, applied to finance documents

The same sixty documents through a naive agent and a guarded one. Attack success falls from 36.7% to 1.7% with no increase in false blocks, and the run exports as a timestamped report tied to a corpus hash.

Analysis with receipts, where the model never sees the source data, applied to financial statements

Statements computed from the profile, then written up live by the model without it ever seeing the statements. Every figure in the commentary is a chip that opens the computation behind it.

Confidence, gate, learn. The simplest honest agent loop, applied to expense categorisation

A month of card transactions categorised, with anything under the confidence gate held for review. An override is written to vendor memory, so the next run covers more from memory and calls the model less.

A monthly finance pack narrated by a model that cannot invent a number, applied to management reporting

Assemble a period, review the draft beside the figures each section was allowed to cite, watch it refuse to issue while gaps are open, waive them with a reason, sign, then verify the archived pack against its hash.

No value the LLM produces can reach a policy decision, applied to invoice approval