Provenance receipt: @cejel/cejel@0.4.9 resolved 0.4.9 (sha256:65b622b42bfbb8376f9955219ef46c8c3601da9432a3e53b16d575dac3f94433). The 0.4.9 npm package carries no provenance attestation, so this digest pins what ran but not where the package came from; see the 0.4.9 changelog entry.
Run environment: Executed published @cejel/cejel@0.4.9 with each declared product name, calibrated default and unchanged corpus commits. Public clones are shallow and blob-filtered; private transparency snapshot retains its previous revision.
Scorer version pin: @cejel/cejel@0.4.9 — reason: published board artifact not re-run for the 0.4.10 or 0.4.11 releases; declared 2026-09-17, re-affirmed 2026-09-25
History
2026-08-18: scores withdrawn and why
What this board claimed. Every row published on 2026-08-18 (Scorer source version:
@cejel/cejel@0.4.3, rubric witan-rubric-v18-prospective-2026-07-25) was
described, in the "How to read this board" section, as "produced through the same sealed
public-scoring entry point used by npx @cejel/cejel ." — a reproducibility
claim: that anyone holding the published package could regenerate these exact numbers from the
pinned commits.
What was established. That claim was false. The scores were computed by this project's
own internal engine copy, reached through an internal leaderboard generator (its location redacted from this record on 2026-09-27; the path is private, the fact is not), never built from —
and independently confirmed to have drifted from — the published @cejel/cejel
package: it was missing detector implementations present downstream in the public source (the
V19–V22 prospective-rubric detectors), and its certificate-rendering layer never received the
0.4.2/0.4.3 label fixes. A single sample repository's score matching the published package's
output was mistaken for proof that the engine itself matched; it did not, and sibling files in the
same internal tree had independently drifted.
What was withdrawn. The scores, ranks, verdicts, and comparable figures for every row,
including the per-repository certificate pages. Corpus membership and this board's methodology
were not withdrawn — the design was never in question, one generation run's provenance was.
Republication condition. Scores return once regenerated by executing the published
@cejel/cejel package itself, end to end, for every row — never an internal copy that
merely resembles it.
Status: MET, 2026-08-19. Every row on this board was produced by shelling out to
npx --yes @cejel/cejel@0.4.4 <path> --out <dir> --quiet against its pinned
commit — a fresh npm resolution, no workspace import, no rubric pin, no internal engine anywhere
in the path. The reproducibility guard that checks this (rebuilt through the same
published-package invocation, rather than the internal comparison that stayed green through the
original defect) confirmed every row matches; a second, independent rescore of one row from a
separate working directory reproduced byte-identical output. Both are recorded in this
regeneration's own record.
How to read this board
Scores run 0.0-4.0 and come from a deterministic rubric over observable repository signals: tests and CI discipline, secret handling, dependency hygiene, audit trail, and governance.
Every score on this board is produced by executing the published `npx @cejel/cejel@0.4.9` package itself, end to end, exactly as an outside caller would: default scoring and rubric settings, plus the declared caller-context `--product-name`; no rubric pin, no workspace import, no internal engine anywhere in the path. The board runs the calibrated public default shown above — full stop; it does not select a prospective rubric. Ordinary public CLI calls (`npx @cejel/cejel@0.4.9 .` with no scoring or rubric flags) get the identical calibrated default this board runs. No private domain collector contributes. Public-repository rows pin an immutable public commit and are independently reproducible from it: cloning that commit and running the same declared invocation a stranger would reproduces the row byte-for-byte. The explicitly labeled private Alfred transparency snapshot is not independently reproducible because its source commit and evidence locations are not public; it is not presented as a public-repository self-score.
Product name and slug are caller context, not scored repository evidence, and are excluded from the byte-comparison claim. For Cejel 0.4.5 and later, this board passes the same declared `--product-name` value when reproducing a row; with that value, the pinned commit, clone recipe, package version, rubric selection, and invocation time held constant, differently named checkout directories emit byte-identical artifacts. If `--product-name` is omitted, Cejel retains its compatibility default of `package.json.name` when usable and the checkout directory name otherwise, so the two identity fields may differ across checkout names.
To reproduce a row exactly, clone at its pinned commit with the same recipe this board uses — a shallow, blob-filtered fetch, not a plain `git clone`: `git init && git remote add origin <url> && git fetch --depth 1 --filter=blob:none --quiet origin <commit> && git checkout --detach --quiet FETCH_HEAD` — then run `npx --yes <name>@<version> <path> --out <dir> --product-name <declared-row-name> --quiet` with the exact package spec in the header above and the row name shown on this board, invoked from a directory that is not itself a checkout of the package being run. A full (non-shallow) clone of the same commit can score differently on history-dependent dimensions (observed on B2/audit-trail evidence) even though the scanned code is byte-identical — the shallow depth and blob filter are the reproduction recipe, not an incidental detail. The invoking directory matters too: running `npx` from inside a local checkout whose own package.json name matches the requested package (for example, a clone of this project itself) can silently resolve to that local checkout regardless of the version pin, rather than fetching the registry version — confirmed live: from such a directory, a pinned `npx --yes <name>@<version> --version` can report an unrelated, older version. Run the command from any directory that is not such a checkout.
Verdict bands are withheld on this board. The previous 3.5 "Verified" cutoff was not calibrated against the comparative population, so attaching that label to corrected aggregation would imply evidence we do not have.
A score reflects observable engineering signals only. It is not a security guarantee, not an audit, and not a judgment of a project's value or its maintainers.
A score reflects only its MEASURED dimensions. The Coverage column shows how many dimensions were actually measured per category (e.g. "code 4/5 · process 1/6"): a dimension that is not applicable to the repository, or that had insufficient data to measure, produced no score and is excluded from the composite rather than counted against the repository. Unmeasured is not good — it is unknown. Coverage counts every rubric dimension, including the two dimensions the repository scanner marks not applicable for every repository.
Rows where fewer than half of the dimensions behind a score were measured are marked "low confidence": low coverage — scored on few signals, less certain. A 4.0 measured from one dimension is weaker evidence than a 3.5 measured from five. A score measured on few dimensions is weaker evidence than a score measured on many, so low-confidence third-party rows are published under "Unranked — insufficient coverage" below rather than ranked against better-evidenced rows.
A repository gets no score, rank, or verdict band when Cejel establishes either structural source absence or zero measurable free-core criteria — see "Unrated — insufficient source or measurable evidence" below. Each row's verdict and reason distinguish those cases. This is different from "low confidence": a low-confidence row is still a real score on few dimensions; an unrated row is not a score at all.
A repository that scores low on a dimension shows its specific findings in the linked evidence report — the findings are the substance, not the verdict.
Overall, Code trust, and Process trust are each repository's rubric-native figures — the exact numbers in its linked evidence report. They are retained for auditability, but the board's order uses a separately named comparative figure.
Comparable score (equal measured criteria) is the board's ordering figure. It excludes two repository-inapplicable process dimensions for every report, then gives every remaining measured criterion equal weight. This applies the pre-statable rule that thin category buckets must not receive disproportionate headline weight; no repository-specific weights or target ranks are used.
The whole corpus membership is published. Third-party repositories form the calibrated public population. Publisher-owned repositories, including Cejel and the private Alfred snapshot, appear only under "Our own code — shown for transparency, not ranked" and receive no public rank. A repository that fails to clone or score appears loudly in its appropriate population and is never silently dropped.
The generator is incremental: a repository with an up-to-date committed evidence report is not re-cloned, so the corpus can grow while every run stays synchronous. Re-scoring everything is a --force flag away.
External repositories are fetched read-only at the immutable commits pinned in corpus.json; none of their code is executed. Re-scoring the same pins is a rubric change. Moving a pin is a separate corpus change.
Evidence map
Third-party repositories only. Overall score is one axis; coverage shows how much of the rubric Cejel could actually measure.
These are each repository’s canonical sub-scores. Low-coverage rows remain visible but are labelled rather than ranked by the chart.
Ranking
This public ranking contains third-party repositories only. It is ordered by "Comparable score (equal measured criteria)": two repository-inapplicable dimensions are excluded uniformly and every remaining measured criterion receives equal weight. Thin buckets therefore cannot take half the headline by construction. Verdict bands are withheld pending calibration. Rows below the coverage floor are excluded from this table. Repositories without enough evidence for any score are excluded here too; see "Unrated — insufficient source or measurable evidence" below — nothing is hidden, only left unordered or unscored.
Publisher-owned repositories are disclosed as transparency snapshots outside the calibrated public population. They receive no rank and no verdict band. Alfred is private and cannot be independently reproduced; Cejel is public and independently inspectable.
Unrated — insufficient source or measurable evidence
We publish repositories we cannot score. A row here means either that Cejel established structural source absence or that a real source tree produced zero measurable free-core criteria. The Verdict and Reason columns distinguish those cases. Neither state is a zero or a low number dressed up as a judgment: Cejel issues no score and no rank when the evidence does not support one.
No free-core rubric criterion produced a measurable signal. Cejel abstains rather than publish a numeric zero for an entirely unmeasured repository. Source coverage: 9 of 250 source-shaped files (3.6%) are criterion-ratable, below the 20% reviewable-source threshold (328 tracked files in total).
Unranked — insufficient coverage
Below the coverage floor: scored on fewer than half of the applicable dimensions, so the score is weaker evidence than a well-covered row. Published in full — same rubric, same numbers — simply not ordered against better-evidenced rows above.