We ran the same test on Laya, an open-weights decision model. Read the first caveat.
Read this first. These rates are on constructed defects built from public pull-request bodies, not on the population of agent reports. And one more thing, specific to this run: Laya has a 512-token window, and 1,098 of the 1,290 question inputs in this run were longer than the room it leaves for the report (257 of 430 in T, 411 of 430 for the probability question in E, and all 430 for the defect-naming question in E). Our preregistration said such states would be refused and never truncated. The primary run did not do that: it used the package's default, which keeps the leading tokens of the state and discards the rest, and it was scored before we noticed. We recorded the departure in a signed erratum and merged the run with this notice on top. So the numbers below describe Laya as packaged, reading the beginning of reports and evidence of this length. They are not the preregistered condition, and under the preregistered rule most cells would have been not evaluable. They are not evidence about what Laya does on short inputs.
Why we ran it
Laya is an open-weights model from Convai Innovations that answers typed questions about a state and returns a probability, with the same request and response shape as the hosted judge we tested on 20 September. It runs locally, so the experiment costs nothing per call. We used the identical 430-case corpus, the identical question wording, and the identical 0.5 threshold, so the two results can be set side by side. We wrote seven predictions down before any call, one of them about Jev's own figure, and committed them in a signed preregistration.
What we did
The frozen Jev corpus: 430 cases from 50 merged public pull requests, each judged text only (T) and with the structured evidence record (E). One pinned model revision (convaiinnovations/laya@55cf4c4e), the ONNX bundle pinned by digest, 860 primary calls. A secondary arm refit one temperature per question on half of the cases and tested on the other half; it made no model call. A third arm, labelled unscored, made 1,098 further calls on exactly the over-window states, using the package's leading-token rule, to show the consequence of the window.
What happened
Report alone, Laya cleared nearly every defective report: premature 47 of 50 (94.0% [83.8, 97.9]), wrong count 47 of 50 (94.0% [83.8, 97.9]), fabricated reference 50 of 50, wrong file 50 of 50. The two perfect cells are printed as counts and withheld as rates under the protocol's perfect-cell rule. Predicted: a judge that sees only the text cannot see the world it describes.
With the evidence attached it still cleared most of them: premature 44 of 50 (88.0% [76.2, 94.4]), fabricated reference 43 of 50 (86.0% [73.8, 93.0]), wrong count 44 of 50 (88.0% [76.2, 94.4]), wrong file 42 of 50 (84.0% [71.5, 91.7]). The prediction that wrong count and wrong file would be at most 20% was contradicted. The evidence lowered the four rates by 6 to 16 points, and none fell below 84%.
The persuasive rewrite raised the probability, where it had lowered it for Jev. Mean probability on defective reports rose by 0.089 [0.064, 0.114] in T and 0.122 [0.085, 0.161] in E, paired over 180 cases. We predicted a rise of at least 0.10 in T and a smaller rise in E; the T interval straddles 0.10 and the rise was larger in E, so the prediction is contradicted.
The probabilities are not calibrated. Expected calibration error 0.702 [0.685, 0.718] in T and 0.642 [0.620, 0.662] in E, n=430 each, against a gate of 0.05, and against 0.458 and 0.420 for Jev on the same cases. We predicted T above 0.10 (supported) and E above Jev's 0.420 (supported).
Refitting the temperature did not repair it. On the held-out half, ECE was 0.428 [0.397, 0.459] in T and 0.419 [0.388, 0.451] in E, against a predicted 0.10 or less. Contradicted. All three fitted temperatures landed on the top of the search grid, 10.00, so the fit is bounded by the grid; a single temperature cannot move a model whose confident answers are wrong in one direction.
Clean reports: false-flagged 2 of 50 in T (4.0% [1.1, 13.5]) and 7 of 50 in E (14.0% [7.0, 26.2]). Both intervals straddle the 10% line, so the prediction is not evaluable.
When asked to name which of five defects a report contained (E only), Laya picked correctly 53 times of 430, 12.3% [9.5, 15.8], below the constant-answer baseline of 23.3%. Jev picked correctly 374 times. The unscored truncation arm applies the same leading-token rule the package applied by default, so on the over-window inputs it repeats what the primary saw; its defect-naming figure matches the primary exactly, 53 of 430, as it should. It confirms the run is deterministic. It is not a separate condition, and it cannot tell us what Laya would do with the whole report.
| Prediction, written before the run | Verdict |
|---|---|
| P1. In T, false-clear on premature and fabricated_reference at least 50% | not evaluable (premature 94.0%; the other cell perfect, withheld) |
| P2. In E, false-clear on wrong_count and scope_mismatch at most 20%, premature higher than both | contradicted |
| P3. Persuasive rewrite raises mean probability by at least 0.10 in T, less in E | contradicted |
| P4. ECE in T exceeds 0.10 | supported |
| P5. False-flag on clean reports below 10% in both conditions | not evaluable (both intervals straddle 0.10) |
| P6. After per-question temperature refit, held-out ECE at most 0.10 in both | contradicted |
| P7. Laya's ECE in E is higher than Jev's 0.420 | supported |
What we take from it
Nothing about Laya's ceiling. On inputs of this length, as packaged, it read the start of the report, cleared nearly all of the false ones, and its probabilities were further from outcomes than Jev's. Two supported predictions, three contradicted, two not evaluable, all published with the same prominence. The method is the point: a preregistered test with a stated threshold, written predictions and a signed record of where the run departed from the plan, on a corpus anyone can re-run at no cost.
Limits, again
Constructed defects, one insertion each, from public pull-request bodies; fifty base cases. One pinned model revision, one wording per question, one threshold. The primary was scored with the package's default truncation, not the preregistered refusal rule; the primary was scored before the erratum that records this. No comparison with any other judge beyond Jev's published figures on the same corpus. Response payloads, call logs and reviewer material are not published.
Records
- Preregistration, signed commit 4c3dee91
- Erratum, recording the truncation departure
- Run approvals
- Result, unedited harness output, with the JSON
- Harness source
- The corpus this run used, unedited: the Jev experiment page and its frozen corpus