Skip to content

lab

PolicyGround

Governed RAG where weak retrieval refuses instead of guessing, applied to a finance policy manual

The same question asked as three roles. Guest and staff are refused with the restricted passages never retrieved, and the controller is answered with every claim carrying a citation.

What it does

Every claim is cited, weak retrieval refuses rather than composing something plausible, sensitivity labels decide what is retrievable at all, and groundedness is gated in CI rather than checked by hand.

The graph is retrieve → assess_sufficiency → (refuse | compose → citation_check). Refusal is a node, not an apology appended to an answer. Which is the difference between a system that declines and one that hedges.

The label filter lives in the retriever protocol rather than in the prompt, so a role that cannot see a document cannot have it summarised at them either. The corpus is a 30-policy synthetic accounting manual written for the project; no real policy text and no accounting-standard text is reproduced.

Screens

The roles demo: one question asked at guest, staff and controller. Guest and staff are refused with passages withheld by the label filter; the controller is answered.
One question, three roles, side by side. Guest and staff are refused with six passages withheld by the label filter; the controller is answered from two restricted sources. The line above the panels is the design decision. The filter runs inside the retriever, so restricted passages are never read rather than read and then hidden. View full size
A refusal: the question scored 0.46 against a 0.80 threshold, with the nearest sections listed under a heading saying they are not an answer.
Refusal is a node, not an apology appended to an answer. It scores the retrieval (0.46 against a 0.80 threshold), says plainly that it will not infer, lists the closest sections under a heading insisting they are not an answer, and logs the question as a possible gap in the manual. View full size
An answered question: each claim carries a numbered citation, with the cited passages shown in a sources panel alongside their sensitivity label and version.
The same machinery when retrieval is sufficient (0.94). Every claim carries a numbered citation and the panel states the rule: uncited claims are removed before render. The prose is blunt because this ran with no model credential. Composition falls back to an extractive stub, while citation, label filtering and refusal stay deterministic and unaffected. View full size

Stack

  • Python
  • LangGraph
  • FastAPI
  • pgvector
  • React

All data in this project is synthetic.