A link is the beginning of a check
NIST’s evaluation-probe project asks three different questions about an AI citation: whether the source supports the claim, whether important context is missing, and whether the evidence is strong enough for the conclusion. Its program page was created in May 2026; this is an explanation of that work, not a new release. [1]
- Match the claim
- Preserve the context
- Limit the conclusion
Make the evidence trail inspectable
The demonstrator evaluates reports against a chosen collection of documents. Language-model judges apply written scoring rules, and their verdicts and reasons are saved beside the report. NIST describes checks that can run during a workflow or afterward. [1]
One trial, three different mistakes
Hypothetical example: Imagine a fictional report: a delivery cart completed a route in a quiet laboratory but was not tested around pedestrians. Saying it failed the route contradicts the source. Saying it succeeded while omitting the quiet-lab condition loses context. Saying it is safe on every busy sidewalk reaches beyond the evidence. These are three different faults, even if every sentence links to the same report.
The checker needs checking too
The repository explicitly says more validation of the probes, scoring rules and model judges is needed before the package can serve as a measurement platform. It is a research demonstrator, not a certificate that cited answers are correct. We reviewed the documentation; we did not run or independently validate the software. [2]
Go a little deeper
Optional reading · about 1 more minute
Use the distinction when you read
Our interpretation: For a consequential answer, try three passes through the cited passage. First compare the actual statement. Then look for conditions the answer left out. Finally ask whether the conclusion demands evidence the source never set out to supply. We find these more useful questions than simply counting the references at the bottom.
A better rewrite
Hypothetical example: For the fictional cart, a restrained sentence would say: “It completed this laboratory route; the report does not establish performance around pedestrians.” That version keeps the useful result and its boundary together. It does not need to dismiss the experiment to avoid overselling it.
Original sources
Attributed synthesis, not original reporting. Examples labeled hypothetical or illustrative are explanatory. Reviewing a source does not independently validate its findings.
- NIST: Building Evaluation Probes into Agentic AI ↗
Program page created May 1 and updated May 5, 2026. Overview, objectives and approach read September 17; describes research goals and testbed, not validated reliability.
- NIST research demonstrator repository ↗
README introduction and workflow reviewed September 17, 2026. Documentation only; no code execution, performance replication or paid API use.
