Releases, corrections, and readings
Dated, sourced, newest first. A release ships, a correction is disclosed, or we read something and say what we think — in the open, and it keeps happening.
-
The release attests what each binary and the image bundle, and makes the certificate read correctly
0.6.1 generates each release SBOM from the build, so it lists the packages each standalone binary and the Docker image actually bundle, and attests it to the artifact it describes. A release with an empty SBOM now fails before anything is published. The certificate’s wording and layout are clearer for a first-time reader; no score, band or verdict moves.
The changelog states who each change affects. The distribution table links the verified release surfaces.
Sources: Changelog: v0.6.1 · v0.6.1 release · Verified distribution table
-
The release adds a GitLab Code Quality export and an includable GitLab CI job
0.6.0 adds
cejel export gitlab-codequality, which writes a report’s located findings as a GitLab Code Quality report, and a GitLab CI template that runs the scan, publishes that report and can fail below a minimum score. No score, band or report field changes. The template has structural tests but has not yet been run on a GitLab instance by us.The changelog states who each change affects. The distribution table links the verified release surfaces.
Sources: Changelog: v0.6.0 · v0.6.0 release · Verified distribution table
-
The release tightens the CI score gate and makes the HTTP MCP server stateless
0.5.0 makes
--min-scorefail when fewer than half of the applicable dimensions were measured, even if the score clears the minimum; scores themselves do not change. The HTTP MCP server now returns certificates and badges with the scan that produced them, and reports move to format 1.4, recording the weight share each metric contributed.The changelog states who each change affects. The distribution table links the verified release surfaces.
Sources: Changelog: v0.5.0 · v0.5.0 release · Verified distribution table
-
The release fixes missed secrets and error handlers, and makes withheld evidence explicit
0.4.11 detects populated PostgreSQL passwords in committed environment files and recognizes more registered Express error-handler idioms. Every certificate now states whether content was withheld and which signals would have read it; the machine report carries the same disclosure.
The changelog records the fixes, paired measurements, and remaining limits. The distribution table links the verified release surfaces.
Sources: Changelog: v0.4.11 · v0.4.11 release · Verified distribution table
-
Jev judge follow-up: the operator post
The operator’s LinkedIn follow-up on the Jev judge test is linked here as its own dated post. The experiment page carries the preregistrations, results, and limits of the primary test and subsequent composition and anchoring arms.
Sources: Read the LinkedIn post · Experiment and follow-up record
-
Judge follow-ups: ask about one claim, then compose the verdict in code
Three preregistered follow-ups tested composition and anchoring on the same frozen corpus and pinned judge. They support a narrower use: ask atomic questions about a claim in front of its evidence, read the label, and compose the verdict in code.
The integrated probability still did not establish a calibrated report-verification decision. The published record includes contradicted predictions, remaining date-comparison weakness, and an erratum caught by the harness before the first call.
-
A preregistered test of a commercial judge model
The published Jev judge experiment compares constructed reports against frozen evidence. Three of five preregistered predictions were contradicted; the page publishes those outcomes alongside the predictions that held, the corpus, harness, and stated limits.
Sources: Primary experiment record
-
0.4.10 repairs npm provenance prospectively
The published 0.4.10 package carries npm attestations from GitHub Actions. The release record verifies identity across the distribution routes; the immutable 0.4.9 package remains disclosed as published without npm provenance.
Sources: Changelog: v0.4.10 · v0.4.10 release
-
The v22 detection-recall measurement: the operator post
The operator’s LinkedIn post on the v22 detection-recall measurement now has its own dated entry. The methodology page provides the measurement’s scope, interval, instrument, and distinction from rubric agreement and fixture-scoped results.
Sources: Read the LinkedIn post · Measurement and scope
-
0.4.9 states the boundary of repository evidence
Certificates describe the repository tree at its pinned revision. Evidence outside that tree is neither seen nor claimed to be absent. The changelog also records paired scoring changes under the same rubric and the historical npm provenance gap.
Sources: Changelog: v0.4.9 · v0.4.9 release
-
Public language, private belief, and what closes the gap
Jacob Coxon, a pretraining researcher, resigned from Anthropic on 2026-09-09 with a public statement that both frontier labs are racing toward self-improving systems.
Our reading, as posted: the argument is about frontier pretraining, a layer we do not work at; the pattern underneath it is one we see everywhere — public language and private belief diverging, and the only thing that has ever closed that gap is evidence a third party can re-run rather than a statement taken on trust. True of a lab’s safety claims, true of a vendor’s software.
Sources: Jacob Coxon’s public statement · Later operator reflection on Coxon’s resignation (LinkedIn)
-
Public language and private belief: the operator post
The operator’s LinkedIn reflection on Jacob Coxon’s resignation is listed separately from the following day’s reading. That reading connects the gap between public claims and private belief to evidence a third party can re-run.
Sources: Read the LinkedIn post · The accompanying reading
-
cejel.dev’s release record is verified across every surface
The site’s release record was completed after the release passed its own currency verifier on all fourteen distribution surfaces: npm with provenance attestations, the GitHub release with five signed binaries, the OCI image, the MCP Registry (now serving 0.4.8), the Homebrew tap, the GitHub Action floating tag, and this site.
The public leaderboard is still scored by 0.4.5 and says so; it is re-scored when the regeneration can be verified end to end.
Sources: Distribution table on /for-engineers/ · cejel verify-release-currency run 34365914100
-
A limit we imposed on ourselves was scored as evidence lost
A design partner ran two versions of Cejel on the same commit and the scores were far apart. Two compounding defects, both ours, both present since 0.4.6: a file over our own 512 KB read limit was mapped by the shape of its filename to criteria that had readable evidence elsewhere, and the resulting abstention was recorded like a file we could not read at all, so it stayed in the composite at zero.
0.4.8 fixes both, words the certificate so the two cases read differently, and records the producing tool version in every report. Disclosed the same day, fixed the same day, credited in the changelog without a name.
Sources: Changelog: v0.4.8 · v0.4.8 release · Later operator post: v22 detection-recall measurement (LinkedIn)
-
0.4.7: findings first
The certificate now shows critical and warning findings above the prose, and every criterion shows the weight actually applied. Presentation only; the machine artifacts are unchanged.
Sources: v0.4.7 release
-
AIR raises $50M for agent add-on vetting
AIR came out of stealth on 2026-09-01 with $10M from Sequoia and $40M from Greenoaks to discover agents inside enterprises, vet the skills and add-ons they load, and enforce a maintained allow-list.
A neighbour, not a competitor: AIR gates what an agent may run at runtime; Cejel produces re-runnable evidence about a codebase at the moment someone has to accept it. Gates consume evidence; the evidence has to come from somewhere the gate does not control.
Sources: TechCrunch, 2026-09-01
-
We withdrew every public score for a day
On 2026-08-18 we found our own leaderboard’s reproducibility claim was false: the published scores had been computed by an internal engine copy that had drifted from the published
@cejel/cejelpackage, not by running the package itself as claimed. Every score, rank, and verdict was withdrawn the same day; corpus membership and methodology were not — the design was never in question, one generation run’s provenance was.Scores returned the next day, once every row was regenerated by shelling out to the published package end to end, with a reproducibility guard now checking exactly that.
Sources: /leaderboard/ withdrawal record
-
A determinism claim about our own tool was wrong
The 0.4.1 release fixed a determinism defect: the checkout path leaked into
report.json, so identical repository content scanned from different directories produced different reports. The claim of byte-identical reproducibility was withdrawn, the fix shipped, and the claim was reinstated with a test guarding it.Sources: Changelog: v0.4.1 · Later operator post: Jev test follow-up (LinkedIn)