Readings and limits
Results and Measurements
This page belongs to the folder of the Verifiable natural-language generation project. It sets out the readings of 28 August 2026 on the current state of the system under study: fact base, linguistic coverage, factual equivalence and determinism. Every figure is accompanied by the table of the same data, and the last section names what these figures do not say.
Measurements of , regenerable on the current state of the system.
1. How these measurements are made
The figures on this page are not estimates, and they are not preserved records either. They are produced by readings that are redone in one command on the current state of the system, and that overwrite the previous ones. The date displayed at the top of the page is that of the reading, not that of the writing: when the system evolves, the measurement is redone and the page changes.
The editorial principle that follows holds in one sentence: the division publishes only what it can re-measure. A result that a reading could not reproduce has no place here, even if it is true. It is also the reason why no numerical comparison with other systems appears on this site: the division measures what it controls, and nothing else.
Three families of readings share the page. The first describes what the system has at its disposal, that is, its base of structured facts. The second measures what it actually produces, language by language and subject by subject. The third submits the outputs to external checks, among them a negative control intended to establish that those checks know how to fail.
Three operational definitions fix the vocabulary of the readings.
Definition 1 (complete document). For a given subject and language, an output returned in full, with no empty section and no production error. The sweep counts a language as covering if and only if that document exists.
Definition 2 (honest omission). The counted absence of a document, or of a fact within one, where the available material does not allow it to be stated correctly. Every omission is published next to the measurement it belongs to: coverage and omissions add up to 1,000 by construction.
Definition 3 (factual equivalence). Two versions of one subject are equivalent if every data value extracted from either, at the published threshold, is found without conflict in the subject’s reference set. Saying less is permitted, saying something else is not.
2. The fact base, by domain
The system works on 3,444 structured fact sheets, spread across 8 domains. The distribution is very uneven, and this inequality explains a large share of the coverage gaps observed further on: a domain with few sheets, or whose sheets are sparse, mechanically produces fewer complete documents.
| Domain | Fact sheets |
|---|---|
| Historical figures | 1,070 |
| Cities and places | 1,003 |
| Living species | 504 |
| Events | 445 |
| Countries | 197 |
| Substances | 156 |
| Inventions | 65 |
| Celestial objects | 4 |
| Total | 3,444 |
3. Languages producing a document
The sweep of 28 August 2026 counts, for a given test subject, how many languages produce a complete document out of the 1,000 languages processed. The remaining languages do not produce a degraded document: they keep silent, and that silence is counted as an honest omission.
| Test subject | Document produced | Honest omissions |
|---|---|---|
| Country (Kenya) | 935 | 65 |
| Living species (lion) | 932 | 68 |
| Substance | 924 | 76 |
| Historical figure | 132 | 868 |
The two columns of figures add up to 1,000: every language that does not produce a document is counted as an honest omission, and none is lost along the way. The length of the documents varies strongly with the subject: the median settles at 77 words for the country subject, at 8 words for the species and at 7 words for the substance. A French example of a country sheet runs to 274 words, at a measured reading level of grade 12.
Publishing the omissions column is not a rhetorical precaution; it is a condition for reading the other columns. A high coverage rate only means something if one knows what became of the rest: without that column, nothing distinguishes a system that keeps silent from a system that approximates. Counting the silences is what makes the assertions readable, and that is why the least covered domain appears in the same figure and the same table as the best endowed.
4. Robustness across subjects
A reading taken on a single subject per domain lends itself to convenient choice. The multi-subject reading rules that risk out by construction: five subjects are retained per domain, chosen in lexicographic order, hence without discretionary intervention. The value carried for each domain is the average number of languages producing a document over those five subjects.
| Domain | Average over 5 subjects |
|---|---|
| Substances | 924 |
| Living species | 908 |
| Countries | 907 |
| Events | 900 |
| Cities and places | 897 |
| Historical figures | 273 |
Five domains out of six hold within a narrow band, which indicates that coverage does not depend on the subject chosen but on the material available. The historical-figures domain confirms its dispersion instead: depending on the density of the sheet, the number of languages producing a document runs from 109 to 902. The limit is therefore the available data, and not the production procedure, which the division says rather than dropping the domain from the reading.
The same multi-subject reading measures the length of the documents produced. Figure 4 carries the thirty medians, one per subject: the position of a cluster tells the richness of the domain, its spread the variance within the domain.
The substance cluster is one repeated dot: a median of 7 words on each of the five subjects, the signature of a domain whose available material is uniform from subject to subject. At the opposite end, the country domain spreads from 19 to 68 median words depending on the subject. Historical figures combine the two readings: coverage there varies from 109 to 902 languages depending on the sheet, but the median length stays narrow, from 11 to 13 words – what gets said varies little; what varies is the number of languages able to say it.
5. Factual equivalence and determinism
The two readings that follow bear on the outputs themselves. The first verifies factual equivalence between languages, in the sense of definition 3: saying less is permitted, saying something else is not. The texts produced were re-read by a verifier independent of the system, which extracts the numerical data values and checks that each is found, without conflict, in the subject’s reference set, at a published threshold: values at or above 10,000, which excludes the years carried by proper names. Out of 935 languages re-read, 415 values were compared, and no conflict was found.
A check that finds nothing has value only if it is capable of finding something. A negative control therefore accompanies the reading: a corrupted value artificially injected into an output was duly detected by the verifier. It is this capacity to fail that makes the previous result readable as a result, and not as an instrument’s silence.
The second reading bears on determinism. 147 generations were repeated in two separate processes, and the 147 SHA-256 fingerprints obtained are identical: no deviation. The reproducibility claimed is therefore not a postulated internal property, but an equality observed from outside, on fingerprints that any third party knows how to compute.
Diagram 1. Inter-process determinism, reading of 28 August 2026: the same 147 generations, replayed in two separate processes, return identical fingerprints.
| Reading | Value |
|---|---|
| Languages re-read by the independent verifier | 935 |
| Numerical values compared | 415 |
| Conflicts found | 0 |
| Generations repeated in two separate processes | 147 |
| Identical SHA-256 fingerprints | 147 |
| Deviations found | 0 |
Property P4
Factual equivalence between languages
The versions of one subject in different languages carry exactly the same facts: saying less is permitted, saying something else is not.
Status: verified from outside the system on the reading of 28 August 2026, negative control included.
6. The extended campaign
The readings above concern one trial subject. The same verification was then replayed at scale: 30 subjects, chosen in lexicographic order as the first 5 of each of the 6 measured domains, re-read across the 1,000 languages. 1,646 numerical values were compared from the outside, and no conflict was found. The negative controls were replayed subject by subject: on the 8 subjects whose reference set carries values above the threshold, 8 injected corruptions, 8 detections.
Diagram 2. The circuit of the extended campaign, 28 August 2026. The judge is external: it reads only the finished outputs. The lower branch is the negative control, which establishes that the comparison can fail.
The verifier’s calibration was measured separately, in both directions. 105 artificial corruptions were injected into real outputs, one altered value per text, across 15 languages and 8 subjects: 105 detections out of 105. Conversely, 113 untouched texts were re-read without a single false alarm. The instrument’s sensitivity and specificity are measured, not assumed.
Diagram 3. Calibration of the verifier, campaign of 28 August 2026. Both error cells are empty, and both are published: an instrument is judged on its two columns.
Two coverage readings complete the campaign: all 197 subjects of the country domain produce a document in each of the 4 witness languages (French, English, Swahili, Arabic), and 25 distinct writing systems were detected in the outputs themselves.
| Reading | Value |
|---|---|
| Subjects re-read (5 per domain, lexicographic order) | 30 |
| Numerical values compared across languages | 1,646 |
| Conflicts found | 0 |
| Artificial corruptions detected | 105 / 105 |
| False alarms on untouched texts | 0 / 113 |
| Country-domain subjects served in all 4 witness languages | 197 / 197 |
| Writing systems detected in the outputs | 25 |
7. The writing systems, in detail
The dominant writing system of each document produced during the campaign was detected from outside, by Unicode ranges, without looking at the procedure that produced it. 25 distinct writing systems emerge. The Latin alphabet dominates, and 21 of the 25 systems are carried by only one or two languages each: the typographic long tail that equal treatment of languages has to serve correctly, down to Tibetan or Odia.
| Writing system | Languages |
|---|---|
| Latin | 895 |
| Cyrillic | 9 |
| Arabic | 6 |
| Devanagari | 3 |
| Ethiopic | 2 |
| Bengali | 2 |
| Hebrew | 2 |
| Han characters | 2 |
| Greek | 1 |
| Armenian | 1 |
| Georgian | 1 |
| Hiragana and katakana | 1 |
| Hangul | 1 |
| Thai | 1 |
| Lao | 1 |
| Burmese | 1 |
| Tibetan | 1 |
| Tamil | 1 |
| Telugu | 1 |
| Kannada | 1 |
| Malayalam | 1 |
| Odia | 1 |
| Gurmukhi | 1 |
| Gujarati | 1 |
| Sinhala | 1 |
8. Named limits
The historical-figures domain is the least covered of all, and the gap is not marginal: 132 languages out of 1,000 for the test subject, 273 on average over five subjects, with a dispersion from 109 to 902 depending on the density of the sheet. This domain is also the one that holds the most sheets, which rules out the explanation by scarcity of subjects and points to the density of the material available in each language.
Document length is low outside the country domain. The medians of 8 and 7 words read for the species and the substance describe brief documents, usable as fact sheets but not yet as complete educational content. The division prefers publishing these medians to keeping them quiet: they say where the work remains to be done.
Finally, all the figures on this page are measured on the current state of the system and regenerated at each evolution. They do not describe a frozen system, and a later reading may move them in either direction. The date at the top of the page is therefore part of the result, not an archive note.
9. Threats to validity
The scope of these readings calls for five reservations, which the division prefers to state itself rather than leave to be discovered.
The equivalence verifier compares only numerical values at or above 10,000, a published threshold that excludes the years carried by proper names. 22 of the 30 subjects of the extended campaign carry no value above that threshold: for those subjects, equivalence rests on the system’s internal mechanical checks, not on this external judge.
Equivalence bears on data values at the published threshold, not on display precision. One version may round a value while saying so – writing “about 1,208,000” where another version writes 1,208,333 –, and this is visible in the demonstrations: the verifier compares each extracted value against the subject’s reference set, which carries both forms, and an announced rounding falls under saying less, never under saying something else.
The trial subjects remain few: one per domain for the sweep, five for robustness. Lexicographic order rules out convenient selection; it does not, by itself, establish that the chosen subjects are representative.
Determinism is established between two separate processes on one machine; reproducibility across distinct machines has not yet been the object of a published reading.
Finally, coverage counts documents produced, never their richness: the median lengths published with the sweep bound that reading, and medians of 7 and 8 words outside the country domain describe fact sheets, not lessons.
Cite this page
The measurements on this page are dated and regenerable: a citation therefore carries the measurement date, never a consultation date alone.
ODERSA Research (research division of ODERSA). “Results and Measurements”, measurements of 28 August 2026. https://research.odersa.org/en/results
In BibTeX:
@misc{odersa_research_measurements_2026_en,
author = {{ODERSA Research}},
title = {Results and Measurements},
organization = {ODERSA},
year = {2026},
howpublished = {\url{https://research.odersa.org/en/results}},
note = {Measurements of 28 August 2026, regenerable}
}
Read next
- Validation Approach – How these figures know how to fail: negative controls, the register of properties, independent references.
- Demonstrations – The outputs these readings measure, exactly as produced, in twelve languages.
- Data – The same indicators, machine-readable: re-read them without taking our word.