Cejel 0.6.1 is out. See what changed →
For teams that accept code they didn't write

Know what you're accepting.

Cejel checks a codebase and hands you a certificate anyone can re-run. Your vendor, your customer and your auditor run the same command and get the same answer.

  • Your code never leaves your machine
  • Free and open source
  • Limits stated, never hidden
Cejel certificate · example
your-org/your-repo @ pinned revision
Pinned
  • What the evidence showsEvery finding points to a file and a line.
  • What it could not establishThe gaps are written down, not left out.
  • Anyone can check itSame command, same revision, same bytes.
$npx @cejel/cejel@0.6.1 .
How it works

A handoff should come with proof, not a promise.

Three steps, and both sides can see every one of them.

1

Pin the delivery

Pick the exact revision you are being asked to accept.

2

See the evidence and the gaps

Findings you can click through, and a plain list of what was not checked.

3

Share it, let them re-run it

Send the certificate. The other side runs one command and compares.

Product demo · 96 seconds

See what the receiving team gets.

The published package scans a pinned public revision, writes its certificate, states its limits, and produces the same report in a second clean checkout.

The result is evidence to decide with, not a guarantee that the software is safe.

Recorded on 0.4.6. The 0.6.1 certificate wording for size-limit skips differs; see the changelog.

Measured, not claimed

We test ourselves the way we test everyone else.

Written down before the test, published after it, including the parts that went against us.

96 in 100 flags hold up

When Cejel flagged an issue on 200 open-source projects it had never seen, 96 out of 100 flags held up when checked.

96.43%, lower bound 94.16%, test fixed in advance.

How we measured it →

We publish our mistakes

In August we found a claim on our own leaderboard was wrong. We pulled every score, said so publicly, and left the notice up.

What happened →

Same version, everywhere

npm, Homebrew, Docker, the GitHub Action and three more, each checked after every release.

See the checks →
Experiments

We test the AI judges other people trust.

AI reviewers and judges now approve other AI's work. We give them test cases where we already know the right answer, and publish how they did.

Jev, by TypeSafe AI20 Sep 2026

A commercial AI judge waved through 9 in 10 false reports.

430 test cases, four minutes, three US cents. Showing it the evidence helped on only one kind of error.

We wrote down five predictions first. Three were wrong, and they are published next to the two that held.
Read the experiment →
Laya, by Convai Innovations29 Sep 2026

An open AI model passed most false reports, even with the evidence in front of it.

Same 430 cases. With the evidence attached it still cleared 84 to 88 in 100 of the false ones.

Read the caveat first: most test inputs were longer than the model's window, so it judged the start of each report.
Read the experiment →

Both tests use constructed cases, so the rates describe those cases, not how often real reports are wrong. All experiments →

Ways to start

Start free. Bring us in when a decision is on the line.

Free tool

Run it yourself

Offline, no account, nothing uploaded. Open source under AGPL-3.0.

Also: pnpm dlx @cejel/cejel@0.6.1 . · bunx @cejel/cejel@0.6.1 .

$npx @cejel/cejel@0.6.1 .
Latest

Releases, corrections and readings, as they happen.

Have a delivery to accept?

Tell us the repository and the decision that depends on it. We agree the scope before anything runs.