A commercial AI judge waved through 9 in 10 false reports.
430 test cases, four minutes, three US cents. Showing it the evidence helped on only one kind of error.
Cejel checks a codebase and hands you a certificate anyone can re-run. Your vendor, your customer and your auditor run the same command and get the same answer.
Three steps, and both sides can see every one of them.
Pick the exact revision you are being asked to accept.
Findings you can click through, and a plain list of what was not checked.
Send the certificate. The other side runs one command and compares.
The published package scans a pinned public revision, writes its certificate, states its limits, and produces the same report in a second clean checkout.
The result is evidence to decide with, not a guarantee that the software is safe.
Recorded on 0.4.6. The 0.6.1 certificate wording for size-limit skips differs; see the changelog.
Written down before the test, published after it, including the parts that went against us.
When Cejel flagged an issue on 200 open-source projects it had never seen, 96 out of 100 flags held up when checked.
96.43%, lower bound 94.16%, test fixed in advance.
How we measured it →In August we found a claim on our own leaderboard was wrong. We pulled every score, said so publicly, and left the notice up.
What happened →npm, Homebrew, Docker, the GitHub Action and three more, each checked after every release.
See the checks →AI reviewers and judges now approve other AI's work. We give them test cases where we already know the right answer, and publish how they did.
430 test cases, four minutes, three US cents. Showing it the evidence helped on only one kind of error.
Same 430 cases. With the evidence attached it still cleared 84 to 88 in 100 of the false ones.
Both tests use constructed cases, so the rates describe those cases, not how often real reports are wrong. All experiments →
Offline, no account, nothing uploaded. Open source under AGPL-3.0.
Also: pnpm dlx @cejel/cejel@0.6.1 . · bunx @cejel/cejel@0.6.1 .
One repository, one decision. We run it with you in your environment and hand you a two-page memo you can give to whoever has to sign.
Enquire → What's includedTell us the repository and the decision that depends on it. We agree the scope before anything runs.