This is an official ODERSA website. Here’s how you know

The official domain

The address of this site ends in odersa.org. Every service the association runs sits on a subdomain of odersa.org and nowhere else. If the address in your browser’s address bar does not end in odersa.org, this site is not ours.

Free, and no account

Everything is open straight away. No sign-up, no account, no password, no subscription, no advertising. Nothing is held back for those who pay, because there is nothing to pay for.

No data collected

This site does not follow you: no tracker, no tracking cookie, no measurement tool built into these pages, and nothing measured on your device. Our host counts requests in aggregate, as any server that answers does: a total, never a profile. You do not have to take our word for it: open your browser’s developer tools, go to the Network tab, and reload the page. You will see the full list of what the site asks for. Everything comes from odersa.org, nothing goes anywhere else.

Free to reuse

The content is published under the CC BY 4.0 licence. You may copy it, translate it, print it and pass it on, for your classes as much as for the people around you, on one condition only: credit ODERSA.

Readings and limits

Results and Measurements

This page belongs to the folder of the Verifiable natural-language generation project. It sets out the readings of 28 August 2026 on the current state of the system under study: fact base, linguistic coverage, factual equivalence and determinism. Every figure is accompanied by the table of the same data, and the last section names what these figures do not say.

Measurements of , regenerable on the current state of the system.

Measurement date: 28 August 2026 Readings regenerable on the current state

1. How these measurements are made

The figures on this page are not estimates, and they are not preserved records either. They are produced by readings that are redone in one command on the current state of the system, and that overwrite the previous ones. The date displayed at the top of the page is that of the reading, not that of the writing: when the system evolves, the measurement is redone and the page changes.

The editorial principle that follows holds in one sentence: the division publishes only what it can re-measure. A result that a reading could not reproduce has no place here, even if it is true. It is also the reason why no numerical comparison with other systems appears on this site: the division measures what it controls, and nothing else.

Three families of readings share the page. The first describes what the system has at its disposal, that is, its base of structured facts. The second measures what it actually produces, language by language and subject by subject. The third submits the outputs to external checks, among them a negative control intended to establish that those checks know how to fail.

Sloping wooden type case filled with lead printing type arranged by compartment, with a composing stick resting in the centre.
Type case, Gutenberg-Museum Mainz, photo dronepicr, CC BY 2.0.

Three operational definitions fix the vocabulary of the readings.

Definition 1 (complete document). For a given subject and language, an output returned in full, with no empty section and no production error. The sweep counts a language as covering if and only if that document exists.

Definition 2 (honest omission). The counted absence of a document, or of a fact within one, where the available material does not allow it to be stated correctly. Every omission is published next to the measurement it belongs to: coverage and omissions add up to 1,000 by construction.

Definition 3 (factual equivalence). Two versions of one subject are equivalent if every data value extracted from either, at the published threshold, is found without conflict in the subject’s reference set. Saying less is permitted, saying something else is not.

2. The fact base, by domain

The system works on 3,444 structured fact sheets, spread across 8 domains. The distribution is very uneven, and this inequality explains a large share of the coverage gaps observed further on: a domain with few sheets, or whose sheets are sparse, mechanically produces fewer complete documents.

Structured fact sheets by domain: historical figures 1,070; cities and places 1,003; living species 504; events 445; countries 197; substances 156; inventions 65; celestial objects 4. historical figures 1,070 cities and places 1,003 living species 504 events 445 countries 197 substances 156 inventions 65 celestial objects 4
Figure 1. Structured fact sheets held by the system, by domain, on 28 August 2026. The same values are given in Table 1.
Table 1. Structured fact sheets by domain, 28 August 2026
Domain Fact sheets
Historical figures1,070
Cities and places1,003
Living species504
Events445
Countries197
Substances156
Inventions65
Celestial objects4
Total3,444

3. Languages producing a document

The sweep of 28 August 2026 counts, for a given test subject, how many languages produce a complete document out of the 1,000 languages processed. The remaining languages do not produce a degraded document: they keep silent, and that silence is counted as an honest omission.

Languages producing a complete document, out of 1,000 processed languages, 28 August 2026 sweep: country (Kenya) 935; species (lion) 932; substance 924; historical figure 132. country (Kenya) 935 / 1,000 species (lion) 932 / 1,000 substance 924 / 1,000 historical figure 132 / 1,000
Figure 2. Languages producing a complete document for one test subject per domain, sweep of 28 August 2026. The same values are given in Table 2, with the column of honest omissions.
Table 2. Sweep of 28 August 2026: languages producing a complete document and honest omissions, out of 1,000 languages
Test subject Document produced Honest omissions
Country (Kenya)93565
Living species (lion)93268
Substance92476
Historical figure132868

The two columns of figures add up to 1,000: every language that does not produce a document is counted as an honest omission, and none is lost along the way. The length of the documents varies strongly with the subject: the median settles at 77 words for the country subject, at 8 words for the species and at 7 words for the substance. A French example of a country sheet runs to 274 words, at a measured reading level of grade 12.

Publishing the omissions column is not a rhetorical precaution; it is a condition for reading the other columns. A high coverage rate only means something if one knows what became of the rest: without that column, nothing distinguishes a system that keeps silent from a system that approximates. Counting the silences is what makes the assertions readable, and that is why the least covered domain appears in the same figure and the same table as the best endowed.

4. Robustness across subjects

A reading taken on a single subject per domain lends itself to convenient choice. The multi-subject reading rules that risk out by construction: five subjects are retained per domain, chosen in lexicographic order, hence without discretionary intervention. The value carried for each domain is the average number of languages producing a document over those five subjects.

Languages producing a document, average over 5 subjects per domain chosen in lexicographic order, out of 1,000 languages, 28 August 2026: substances 924; living species 908; countries 907; events 900; cities and places 897; historical figures 273. substances 924 / 1,000 living species 908 / 1,000 countries 907 / 1,000 events 900 / 1,000 cities and places 897 / 1,000 historical figures 273 / 1,000
Figure 3. Languages producing a document, average over 5 subjects per domain chosen in lexicographic order, 28 August 2026. The same values are given in Table 3.
Table 3. Multi-subject robustness, 28 August 2026: average number of languages producing a document over 5 subjects per domain, out of 1,000 languages
Domain Average over 5 subjects
Substances924
Living species908
Countries907
Events900
Cities and places897
Historical figures273

Five domains out of six hold within a narrow band, which indicates that coverage does not depend on the subject chosen but on the material available. The historical-figures domain confirms its dispersion instead: depending on the density of the sheet, the number of languages producing a document runs from 109 to 902. The limit is therefore the available data, and not the production procedure, which the division says rather than dropping the domain from the reading.

The same multi-subject reading measures the length of the documents produced. Figure 4 carries the thirty medians, one per subject: the position of a cluster tells the richness of the domain, its spread the variance within the domain.

Median document length in words, five subjects per domain, campaign of 28 August 2026: countries from 19 to 68; events from 18 to 30; cities and places from 8 to 16; living species from 11 to 14; historical figures from 11 to 13; substances 7 on each of the five subjects. countries 19–68 events 18–30 cities and places 8–16 living species 11–14 historical figures 11–13 substances 7
Figure 4. Median document length, in words, for each of the thirty campaign subjects (five per domain, lexicographic order), 28 August 2026. Each dot is one subject; the axis runs from 0 to 70 words.

The substance cluster is one repeated dot: a median of 7 words on each of the five subjects, the signature of a domain whose available material is uniform from subject to subject. At the opposite end, the country domain spreads from 19 to 68 median words depending on the subject. Historical figures combine the two readings: coverage there varies from 109 to 902 languages depending on the sheet, but the median length stays narrow, from 11 to 13 words – what gets said varies little; what varies is the number of languages able to say it.

5. Factual equivalence and determinism

The two readings that follow bear on the outputs themselves. The first verifies factual equivalence between languages, in the sense of definition 3: saying less is permitted, saying something else is not. The texts produced were re-read by a verifier independent of the system, which extracts the numerical data values and checks that each is found, without conflict, in the subject’s reference set, at a published threshold: values at or above 10,000, which excludes the years carried by proper names. Out of 935 languages re-read, 415 values were compared, and no conflict was found.

A check that finds nothing has value only if it is capable of finding something. A negative control therefore accompanies the reading: a corrupted value artificially injected into an output was duly detected by the verifier. It is this capacity to fail that makes the previous result readable as a result, and not as an instrument’s silence.

The second reading bears on determinism. 147 generations were repeated in two separate processes, and the 147 SHA-256 fingerprints obtained are identical: no deviation. The reproducibility claimed is therefore not a postulated internal property, but an equality observed from outside, on fingerprints that any third party knows how to compute.

Process A Process B SHA-256 fingerprints 147 generations the same 147 generations computable by any third party 147 / 147 identical 0 deviations

Diagram 1. Inter-process determinism, reading of 28 August 2026: the same 147 generations, replayed in two separate processes, return identical fingerprints.

Table 4. Factual equivalence and determinism, 28 August 2026
Reading Value
Languages re-read by the independent verifier935
Numerical values compared415
Conflicts found0
Generations repeated in two separate processes147
Identical SHA-256 fingerprints147
Deviations found0

Property P4

Factual equivalence between languages

The versions of one subject in different languages carry exactly the same facts: saying less is permitted, saying something else is not.

Status: verified from outside the system on the reading of 28 August 2026, negative control included.

6. The extended campaign

The readings above concern one trial subject. The same verification was then replayed at scale: 30 subjects, chosen in lexicographic order as the first 5 of each of the 6 measured domains, re-read across the 1,000 languages. 1,646 numerical values were compared from the outside, and no conflict was found. The negative controls were replayed subject by subject: on the 8 subjects whose reference set carries values above the threshold, 8 injected corruptions, 8 detections.

30 subjects Outputs produced Values extracted Comparison 8 corruptions lexicographic order 30 subjects, 1,000 languages threshold: 10,000 and above across languages injected on purpose 1,646 values · 0 conflicts 8 / 8 rejected

Diagram 2. The circuit of the extended campaign, 28 August 2026. The judge is external: it reads only the finished outputs. The lower branch is the negative control, which establishes that the comparison can fail.

The verifier’s calibration was measured separately, in both directions. 105 artificial corruptions were injected into real outputs, one altered value per text, across 15 languages and 8 subjects: 105 detections out of 105. Conversely, 113 untouched texts were re-read without a single false alarm. The instrument’s sensitivity and specificity are measured, not assumed.

Verdict of the verifier alarm raised silence Reality corrupted text (105) untouched text (113) 105 / 105 detections 0 missed corruption 0 false alarms 113 / 113 re-read without alarm One altered value per text, 15 languages, 8 subjects.

Diagram 3. Calibration of the verifier, campaign of 28 August 2026. Both error cells are empty, and both are published: an instrument is judged on its two columns.

Two coverage readings complete the campaign: all 197 subjects of the country domain produce a document in each of the 4 witness languages (French, English, Swahili, Arabic), and 25 distinct writing systems were detected in the outputs themselves.

Table 5. Extended campaign, 28 August 2026
Reading Value
Subjects re-read (5 per domain, lexicographic order)30
Numerical values compared across languages1,646
Conflicts found0
Artificial corruptions detected105 / 105
False alarms on untouched texts0 / 113
Country-domain subjects served in all 4 witness languages197 / 197
Writing systems detected in the outputs25

7. The writing systems, in detail

The dominant writing system of each document produced during the campaign was detected from outside, by Unicode ranges, without looking at the procedure that produced it. 25 distinct writing systems emerge. The Latin alphabet dominates, and 21 of the 25 systems are carried by only one or two languages each: the typographic long tail that equal treatment of languages has to serve correctly, down to Tibetan or Odia.

Table 6. Writing systems detected in the campaign outputs, by number of languages, 28 August 2026
Writing system Languages
Latin895
Cyrillic9
Arabic6
Devanagari3
Ethiopic2
Bengali2
Hebrew2
Han characters2
Greek1
Armenian1
Georgian1
Hiragana and katakana1
Hangul1
Thai1
Lao1
Burmese1
Tibetan1
Tamil1
Telugu1
Kannada1
Malayalam1
Odia1
Gurmukhi1
Gujarati1
Sinhala1

8. Named limits

The historical-figures domain is the least covered of all, and the gap is not marginal: 132 languages out of 1,000 for the test subject, 273 on average over five subjects, with a dispersion from 109 to 902 depending on the density of the sheet. This domain is also the one that holds the most sheets, which rules out the explanation by scarcity of subjects and points to the density of the material available in each language.

Document length is low outside the country domain. The medians of 8 and 7 words read for the species and the substance describe brief documents, usable as fact sheets but not yet as complete educational content. The division prefers publishing these medians to keeping them quiet: they say where the work remains to be done.

Finally, all the figures on this page are measured on the current state of the system and regenerated at each evolution. They do not describe a frozen system, and a later reading may move them in either direction. The date at the top of the page is therefore part of the result, not an archive note.

9. Threats to validity

The scope of these readings calls for five reservations, which the division prefers to state itself rather than leave to be discovered.

The equivalence verifier compares only numerical values at or above 10,000, a published threshold that excludes the years carried by proper names. 22 of the 30 subjects of the extended campaign carry no value above that threshold: for those subjects, equivalence rests on the system’s internal mechanical checks, not on this external judge.

Equivalence bears on data values at the published threshold, not on display precision. One version may round a value while saying so – writing “about 1,208,000” where another version writes 1,208,333 –, and this is visible in the demonstrations: the verifier compares each extracted value against the subject’s reference set, which carries both forms, and an announced rounding falls under saying less, never under saying something else.

The trial subjects remain few: one per domain for the sweep, five for robustness. Lexicographic order rules out convenient selection; it does not, by itself, establish that the chosen subjects are representative.

Determinism is established between two separate processes on one machine; reproducibility across distinct machines has not yet been the object of a published reading.

Finally, coverage counts documents produced, never their richness: the median lengths published with the sweep bound that reading, and medians of 7 and 8 words outside the country domain describe fact sheets, not lessons.

Cite this page

The measurements on this page are dated and regenerable: a citation therefore carries the measurement date, never a consultation date alone.

ODERSA Research (research division of ODERSA). “Results and Measurements”, measurements of 28 August 2026. https://research.odersa.org/en/results

In BibTeX:

@misc{odersa_research_measurements_2026_en,
  author       = {{ODERSA Research}},
  title        = {Results and Measurements},
  organization = {ODERSA},
  year         = {2026},
  howpublished = {\url{https://research.odersa.org/en/results}},
  note         = {Measurements of 28 August 2026, regenerable}
}

Read next

Keyboard shortcuts

TabMove from link to link
EnterOpen the link, or expand and collapse a question
?Open this help
EscClose the panel or this help