Tests we ran. Results we published.
We write each test and our predictions down before it runs. Then we publish what happened, including where we were wrong.
Predictions firstFrozen before the first call, so we can't move the goalposts.
Known answersBuilt test cases, so a rate describes those cases, not the real world.
Corrections on the recordAny change is written down and signed, and shown on the result.
6 Oct 2026Jev and Dunstana model judge and a grammar · our interest stated
A model judge and a grammar with no model failed at the same step.
On the same 430 constructed reports, both missed false claims they did not find in the prose. Once a claim was isolated, both mostly checked it; the page states the exception.
Barg Labs builds the grammarboth zeros predicted before the run
Read it →
29 Sep 2026LayaConvai Innovations · open model
An open AI model passed most false reports, even with the evidence in front of it.
Same 430 cases as Jev. With the evidence attached it still cleared 84 to 88 in 100 of the false ones.
Caveat first: most inputs were longer than the model's window, so it judged the start of each report.
Read it →
20 Sep 2026JevTypeSafe AI · commercial judge
A commercial AI judge waved through 9 in 10 false reports.
430 test cases, four minutes, three US cents. Showing it the evidence helped on only one kind of error.
5 predictions written first3 were wrong, published anyway
Read it →
Our own toolCejelBarg Labs · tested on itself
We ran the same discipline on Cejel: 96 in 100 of its flags held up.
200 open-source projects it had never seen, test fixed in advance. 96.43%, lower bound 94.16%.
How we measured it →