Your CI says green. It does not say what it looked at. For code written by people that was a tolerable omission. For work produced by agents it is the whole decision.
Where the bottleneck went
Generation got cheap. Accepting responsibility did not.
An agent can open forty pull requests, run ten thousand evaluations, or produce a research result overnight. None of that is the constraint any more. The constraint is the person who has to put their name on it.
That person is not short of output. They are short of grounds. They need to know what was checked, by what, against which version of the thing, and what nobody got to. Every tool in the pipeline is built to answer the first half of that question and silent on the second.
“We looked and found nothing” and “we did not look” arrive at a reviewer as the same green tick.
Why a score is the wrong answer
One number standing in for a judgement nobody made
The obvious response to too much machine output is to compress it. Score the run, show the score, gate on the score. We think that is precisely backwards, for a reason that has nothing to do with how good the scoring is.
Two million records is not reviewable. A score over two million records is not reviewable either. It is a number that looks like a finding, produced by an aggregation nobody inspected, and it has no place to put the most important fact about itself — which parts of the input it never reached.
A score can be 94. It cannot be “94, and we did not examine the execution path, and one contributor supplied 77% of the supporting evidence.” Those clauses are where the decision actually lives.
What we built instead
Not a verifier. The argument a verifier leaves behind.
release-gate does not generate results and it does not verify them. Agents generate, tools verify, and the thing nobody was building is the layer between: assembling the argument those two leave behind, then saying whether it is sound enough to put to a person.
You give it a run. An OpenTelemetry trace, a Langfuse export, a promptfoo result, a LangGraph export — nine formats, detected from the content rather than declared, because a flag you have to pass is a flag you can pass wrongly. There is no config file, and nothing is read from the filesystem: a gate whose verdict depends on which directory it ran from is a gate whose verdict cannot be reproduced.
Three things come back. A bounded list of what needs a human. What would close each item. And the part this essay is about.
The coverage ledger
Every verdict carries what it did not assess
Not a footnote. A block of the report, printed for every dimension, on every run:
WHAT WAS ASSESSED
[ assessed] input_integrity: the input file was hashed by release-gate
[ assessed] record_mapping: 6 of 6 record(s) mapped; 0 skipped
[NOT_ASSESSED] execution_reconstruction: no execution graph was reconstructed
[NOT_ASSESSED] replication: no verification attempt names a target
[NOT_ASSESSED] adversarial_review: no verifier set out to disprove this
NOT_ASSESSED is a first-class answer, kept apart from we looked
and found nothing everywhere in the engine. So is UNKNOWN, and so
is a refuted check as against an unresolved one. Collapsing those is not a
simplification. It is how a gap comes to read as a clean bill of health.
On the largest demonstration that ships with it, 2,287,133 records from 10,254 producers reduce to 13 items a person reads, 7 of which cannot be dropped. The reduction is not summarisation. Criticality comes from reachability in the claim graph — a claim one agent emitted once is exactly as load-bearing as one four hundred agents discussed.
What falls out of taking it seriously
Four properties, none of them comfortable
Once coverage is something you have to state rather than imply, several conveniences stop being available. Each of these is a property in the code, not a caveat in the docs:
- Agreement is not corroboration. Nine agents restating one preprint are one piece of evidence wearing nine hats. Lineage concentration is measured: on the worked example, 76.9% of the support collapses to a single root.
- Finding no counterexample bounds the search, not the claim. Ten million sampled instances with nothing found is a statement about ten million instances.
- A shorter list is not a better case. Thirteen items instead of forty can equally mean detection got worse. The engine will not report a reduction as an improvement.
- Scale is not confidence. Ten thousand workers agreeing is ten thousand workers agreeing. It is not evidence about the world.
A failed branch is evidence. A contradiction is closed by evidence that answers it — never by a different branch succeeding.
What it will not tell you
The claims we removed on purpose
Not that a release is safe. Not that a result is guaranteed correct. Not that the gate is unhackable — twenty attacks run against the engine itself and one is recorded NOT_DEFENDED, with what bounds it instead, because a security tool that grades its own worst case is not one. Not that a case is uncontested, that hallucinations are solved, or that an expert has been replaced.
Each of those refusals is a property that cannot return anything else, and each is named in the repository alongside the line that enforces it. A verdict authorises an action against an exact state. It does not certify that a machine was right, and conflating the two is the failure this whole exercise exists to avoid.
The takeaway
The industry is converging on a comfortable story: agents will produce, evaluators will grade, and a number will tell you when to ship. The number is the problem. It is the one artifact in the chain with nowhere to record its own blind spots, handed to the one person who needs them most.
We think the honest unit is not a score but a case: what is being claimed, what it rests on, who checked it and against which version, what contradicts it, and what nobody reached. That is more work to produce and more work to read. It is also the only form in which a person can reasonably accept responsibility for something they did not do themselves.
Ask any tool in your pipeline what it did not check. Most of them have no way to answer.
Submit a run, get the case
No config, no YAML, no account. Nine input formats, detected from the content. One PROMOTE / HOLD / BLOCK, with the coverage it was reached under.
# same engine as the browser demo
pip install release-gate
release-gate assure your-run.jsonl
More from Research
Perfect Code, Imperfect Results — for fifty years correct code meant correct software. Agents void that contract.
The State of Agent Code Safety — we scanned 30+ agent frameworks for the risk SAST cannot see, then threw most of the findings out.