Research note · 2026-01
A verifier that knows how to fail
The measurement campaign of 28 August 2026 rests on an external verifier that compares, across languages, the values carried by the documents produced. This note explains why the probative value of that arrangement hangs on one precise property: the instrument knows how to fail, and that capacity for failure is itself measured.
Measurements of – regenerable. This note comments on readings published on the Results and Measurements page; it introduces no new figure.
1. An instrument’s silence is not a result
A check that finds no error can mean two things: that the material is sound, or that the instrument sees nothing. Silence alone does not tell them apart. A verifier that returned “no conflict” on everything submitted to it – accurate texts and false ones alike – would produce exactly the same reading as a reliable verifier applied to an accurate system, and it would prove nothing.
That is why the division publishes no bare negative result. The “0 conflicts” of the 28 August 2026 campaign only means something because another reading, published at the same rank, establishes that the instrument would have sounded if a conflict had existed. That second reading has a name on this site: the negative control – material the verifier must reject is deliberately submitted to it, and its rejection is observed.
2. Calibration, in both directions
An instrument is judged by its two error columns, and both must be measured. The first is the missed corruption: the instrument stays silent in front of a false text. The second is the false alarm: the instrument cries out in front of a sound one. A verifier is only worth something if both columns are empty, and a reading only shows that if it went and tried to fill them on purpose.
The campaign of 28 August 2026 measures both directions. On the sensitivity side, 105 artificial corruptions were injected into real outputs – one altered value per text, across 15 languages and 8 subjects – and all 105 were detected. On the specificity side, 113 untouched texts were re-read without a single false alarm being raised. Both error cells are empty, and both are published: the first says the instrument sees, the second that it does not cry out for nothing.
3. What calibration makes readable
Once the instrument is calibrated, its silence becomes a result. The same verifier re-read the extended campaign: 30 subjects – the first five in lexicographic order in each of the six measured domains, a selection nobody can steer – across the 1,000 processed languages. 1,646 values were compared from one language to another, and no conflict was found. The negative controls were replayed subject by subject: on the 8 subjects whose reference set carries values above the threshold, 8 corruptions injected, 8 detections.
The initial reading, on a single trial subject – 935 languages re-read, 415 values compared, no conflict – then takes its proper place: that of a sample which the larger scale has confirmed. And the phrase that accompanies factual equivalence on this site – saying less is allowed, saying otherwise is not – stops being a motto: it is the property this external judge puts to the test, text by text, language by language.
4. The published threshold, and what it does not cover
The instrument has a boundary, and it is published with it. The verifier only compares numerical values greater than or equal to 10,000 – a threshold that sets aside the years carried by proper names, which would make any lower threshold misleading. The consequence is stated rather than discovered: 22 of the campaign’s 30 subjects carry no value above that threshold, and their equivalence falls to the system’s internal mechanical checks, not to this external judge. A note that kept quiet about that boundary would make the instrument say more than it measures.
The same logic of externality carries the determinism reading: 147 generations replayed in two separate processes return 147 identical SHA-256 fingerprints. The equality is not postulated from the inside; it is observed on fingerprints any third party knows how to recompute.
5. The rule that follows
The editorial rule of this site holds in one sentence: nothing is published that a reading could not reproduce. This note states its corollary: nothing is concluded from a check that does not know how to fail. The readings commented on here, their figures and their limits, can be read on the Results and Measurements page – the “extended campaign” and “threats to validity” sections – and the approach that frames them, on the Validation Approach page.
6. Citing this note
ODERSA Research (research division of ODERSA). “Note 2026-01 – A verifier that knows how to fail”, measurements of 28 August 2026. https://research.odersa.org/en/notes/note-2026-01 · DOI: 10.5281/zenodo.22159029
In BibTeX:
@misc{odersa_research_note_2026_01_en,
author = {{ODERSA Research}},
title = {Note 2026-01 -- A verifier that knows how to fail},
organization = {ODERSA},
year = {2026},
howpublished = {\url{https://research.odersa.org/en/notes/note-2026-01}},
doi = {10.5281/zenodo.22159029},
note = {Measurements of 28 August 2026, regenerable}
}
Signature: ODERSA Research · CC BY 4.0 licence.