Skip to content

lab

FinXPIA

A prompt-injection corpus with benign twins, so it measures discrimination not blocking, applied to finance documents

The same sixty documents through a naive agent and a guarded one. Attack success falls from 36.7% to 1.7% with no increase in false blocks, and the run exports as a timestamped report tied to a corpus hash.

What it does

Around 60 finance-document prompt-injection cases and 60 benign twins, shipped as both a Promptfoo dataset and a PyRIT dataset with a compliance-ready report dashboard. The twins are the point: a corpus that only contains attacks measures blocking, not discrimination.

Defensive tooling, and declawed deliberately. Every case instantiates an already-public documented pattern. There is no novel attack research here. Exfiltration destinations use only reserved unroutable domains and structurally invalid IBANs, and the CSV vector uses the formula-injection shape wrapped in an inert text function. Both are enforced by tests rather than by convention.

Screens

The FinXPIA report dashboard summarising attack success rate, false-block rate and per-vector results.
The compliance report. Attack success and false-block rate side by side, because a corpus that only measures blocking tells you half the story. View full size
A heatmap of attack success across injection vectors and document types.
Per-vector results across document types, with the benign twins scored alongside. View full size

Stack

  • Python
  • Promptfoo
  • PyRIT

All data in this project is synthetic.