The Evidence Problem

Why None Of This Can Currently Be Tested Every statistical claim about this manuscript rests on a transcription that two people hold and nobody else can inspect. The page images are free to everyone. We measured whether they are enough.

This is the only page on this site carrying a measurement of our own, and the measurement is a negative one. It is here because the alternative was to build a corpus we could not validate and publish statistics computed on it, which is precisely how work of this kind goes wrong.

I. There is no public transcription

Every statistical statement about a script depends on a transcription: a machine-readable record of which sign appears where. For the Rohonc Codex no such thing is publicly available. Each route that exists, checked:

RouteState
The one open transcription projectDead. The page returns 404. Its author's own note says he had got a long way and that life intervened, and it was never finished.
Delia Huegel's long-running studySubstantial interpretive work on the illustrations. Not a transcription dataset.
Király and TokaiThey necessarily hold a machine-readable transcription. The journal lists no supplementary data and it has not been published.
Page imagesAvailable and public domain. 227 scans on the Internet Archive under CC0. Eight pages at higher resolution from Hamburg, 2015.

So the images are free to anyone and the data is held by two people. This applies to claims in favour of the codex and against it equally, and it applies to the 2018 reading as much as to the three that failed. A result that cannot be independently checked is not being doubted when someone says so. It is being described.

II. What the public scans can and cannot support

The obvious response is to build the transcription from the images. We tested whether that is realistic.

The 227 public scans are two-page spreads, roughly 915 by 551 pixels, implying about 454 page images against a known 448. Working from local-mean binarisation, because a global threshold drowns in the ink showing through from the reverse of each leaf:

MeasurementResult
Size of a glyph-sized connected componentmedian 12 × 14 pixels, 90th percentile 29 × 22
Components per spreadmedian 443
Rows per page, spreads left wholemedian 6, against a published 9 to 14
Rows per page, spreads split at the guttermedian 10, mean 10.2
Pages landing inside the published 9 to 14 band52 per cent

Splitting the spreads moved the row count into the published band, which says the approach is sound in principle. Fifty-two per cent per page is not sound enough to build a corpus on. Some of the low counts are certainly correct rather than mistaken, since roughly 87 pages carry illustrations and genuinely hold little text, and that has not been separated out.

The blocker is not segmentation. It is validation.

Suppose the segmentation were perfect. It would give the location of every glyph and the identity of none. At twelve pixels across there is very little information per sign with which to separate several hundred classes, and, decisively, there is no ground truth against which a clustering could be checked. A transcription built this way could not be validated, and statistics computed on an unvalidated transcription would look exactly like a result while being an artefact of the clustering. So the corpus was not built.

III. What would change this

The transcription. Király and Tokai have one. Published, or shared with researchers, it would let their reading be tested independently by anyone, which is what their result deserves and has not had. It would also make every conventional structural measurement possible in an afternoon, and it would settle the sign inventory question that has been open since 1892.

Higher-resolution images. Eight pages exist at good resolution. At three to four times the linear resolution of the public scans, automatic sign clustering moves from marginal to genuinely promising, and the validation problem becomes tractable because a human can check a sample against legible images.

Until one of those arrives there is no Rohonc corpus and therefore no Rohonc result. The constraint is access to data, not analysis. That is an unglamorous conclusion and it is the honest one.

IV. What this site does not claim