Reading, not checking. A model judge and a grammar with no model in it failed at the same step.
Our interest, first. Barg Labs builds Dunstan, a check that holds an agent's completion report to the repository's record without a model, and so has an interest in where model judges fall short. Read what follows with that in mind. The judge figures below are copied from our published Jev experiment record and describe constructed defects built from public pull-request bodies, not the population of agent reports.
The finding
Two very different readers were given agent completion reports that each carried one false claim. One was a hosted model judge. The other was a deterministic extractor: a written grammar, no model. Both failed at the same step. Where they failed, they failed at finding the claim in the prose; where the claim was put in front of them, they checked it against the record, with one exception for the judge that we state below.
The model judge
TypeSafe AI's Jev, pinned at one version, was asked whether each report was accurate. Given the report alone, it cleared nine tenths of the false ones. Given the report with the structured evidence record from the GitHub API, it still cleared most: 84.0% of the wrong-count reports and 92.0% of the wrong-file reports. In 174 of the 380 defective calls it named the defect correctly and, in the same response, gave a probability of 0.5 or more that the report matched the evidence.
Handed the claim, it found the fault. In a follow-up on the wrong-file cases, a replication of the original rendering cleared 94.0%. The same cases, with the report truncated to the paragraph that names the false path and the evidence reduced to the changed-file list, cleared 2.0% [0.4, 10.5], 1 of 50. That arm changed two things at once; reducing the evidence alone gave 66.0% [52.2, 77.6], so truncating the report carries most of the effect, and the split between the two is not measured.
Asked one literal question, it answered it. Re-scoring the frozen call log by the judge's own answer to “which defect does this report contain”, in place of “does the report match the evidence”, took the false-clear rate on wrong count from 84.0% to 1.2% [0.2, 6.6] (n=82), and on wrong file from 92.0% to 1.0% [0.2, 5.6] (n=98). No new call was made.
The exception. The premature completion claim, a date comparison, stayed at 45.0% [35.6, 54.8] (n=100) under the same re-scoring. That is a checking failure, at a step the vendor documents as a weakness, and it is why this page does not say the judge never failed at checking.
Every figure in this section is on the Jev experiment record, with its preregistration, errata and unedited results. Intervals are two-sided 95% Wilson.
The deterministic extractor
Dunstan's advisory grammar reads a report's prose and binds each path, count, reference or timestamp to a clause with an asserting verb, using written rules and no model. Whatever it binds, it compares with the record using the same deterministic checks that Dunstan applies to a declared claim. Tokens it cannot bind are dropped.
On the same 430 constructed reports as the judge, extractor 0.1.2 flagged every planted wrong file (98 of 98) and every planted wrong count (82 of 82). It flagged none of the premature completion claims (0 of 100) and none of the fabricated references (0 of 100). That is 180 of 380 overall, 47.4% [42.4, 52.4], short of the 50% bar set before the run. Both zeros were predicted before the run: a colon ends the clause before “merged”, and a “#N” makes the clause about another pull request, so neither sentence shape is read. Of the 215 disagreements it raised across all 430 reports, 180 were the planted claims (precision 83.7% [78.2, 88.1]). On the 50 clean reports it raised one.
These figures are for extractor 0.1.2 and are published, with their method, in Dunstan’s measured figures at tag v0.1.2. They are kept apart from Dunstan’s figures on real pull requests and are never carried in a Dunstan record.
Addendum, 7 October 2026: extractor 0.1.3. On the same 430 reports, extractor 0.1.3 flagged every planted wrong file (98 of 98), every planted wrong count (82 of 82) and every fabricated reference (100 of 100, up from 0), and none of the premature completion claims (0 of 100). That is 280 of 380 overall, 73.7% [69.0, 77.9], which meets the 50% bar. Of the 315 disagreements it raised, 280 were the planted claims (precision 88.9% [84.9, 91.9]). On the 50 clean reports it raised one. All 100 misses are one sentence shape: every premature plant opens “Status as of <time>: …”, the template in the Jev preregistration, which 0.1.3’s merge-time rule does not read. The zero was predicted before the run. These figures are for extractor 0.1.3 and are published in Dunstan’s measured figures. The reason stands, from the grammar’s side: it checks what it reads, and what it does not read is not there to check.
The failure is the same shape as the judge's, from the opposite side. The judge reads everything and does not reliably isolate the claim. The grammar isolates a claim exactly when the sentence matches a rule, and when it does not match, the claim is not there to check.
On public pull requests: two preregistered studies
The figures above are on constructed reports. To see how the reader does on real ones, we ran two studies on public pull requests written by coding agents (GitHub Copilot's coding agent, Claude, Devin, Google Jules), each preregistered in public before any pull request was selected, with every label published. The operator labelled; the reader and the checks are Dunstan's, with no model.
Study 1, 300 pull requests. Most of what the reader proposed was a real claim: 87% to 92% of sampled advisories. Of the checkable ones, 12.7% to 15.0% were false accusations, almost all of them a file named in the prose for some other reason and read as “I changed this file”. About a third of its disagreement notes marked a claim the record genuinely contradicts (14 of 45). Grammar fixes written from half of the sample did nothing on the other half: the preregistered test refuted them, and they were reverted. Result.
Study 2, 150 new pull requests. A check that asks the repository whether a named file exists lowered false accusations from 9 to 6 of about 92 checkable advisories and lost no genuine catch, but did not earn the “mismatch” wording, and the idea behind it was refuted: 4 of 7 misread file names were files that exist. 19 of the 22 genuine false claims were in one pull request, so shares of disagreement notes in that sample describe it more than a typical one. Result.
The same shape as above, on real reports: the reader fails at finding what a report claims about files, and no grammar we wrote generalised. Where it checks something exact, such as an issue a pull request says it closes, its disagreements were mostly real. A disagreement note is still shown as “possible disagreement, unverified”, and Dunstan’s gate checks only what an agent declares.
What follows for design
If the failure is in finding the claims, the remedy is to stop asking a reader to find them. Check claims one at a time against the record, and have the agent that wrote the report declare them, because it knows what it claimed and it is still there before handback. The Jev record reaches the same advice from the judge's side: ask atomic questions about one claim at a time, in front of the evidence, and compose the verdict in code, which is also the vendor's own design guidance.
Dunstan is built that way. An agent writes a declared handback block: files changed, tests run, references closed, head commit, merged at. Dunstan checks each declared claim against the repository's record and writes a verdict anyone can re-run. dunstan template fills the parts of the block that local git can answer. dunstan suggest runs the extractor over the prose outside the block and lists each claim it proposes the block does not declare, for the agent to declare or reword; it is never a verdict. The checker, its specification and the demo records are public at BargLabs/dunstan, tag v0.1.2.
Limits
The judge figures are on constructed defects, one inserted false sentence per report, from fifty public pull-request bodies; the limits stated on the Jev record apply here unchanged. A zero is never evidence that a subject is clean, and no rate on this page estimates how often a real report is wrong. The extractor figures are on the same constructed reports and are bound to the extractor version they name: 0.1.2, and, in the addendum, 0.1.3. Dunstan’s figures on real pull requests, a different corpus for a different question, are in its measured figures and are not repeated here. Barg Labs builds the extractor and the declared-block check this page recommends.
Records
- The Jev experiment record, with its preregistration, result, composition result (Arm A) and anchoring result
- Dunstan's claim format specification and the extractor’s figures on constructed reports (tag
v0.1.2), and for 0.1.3 here - The two studies on public pull requests: study 1 and study 2, with their preregistrations, amendment, labels and figures