Why None Of This Can Currently Be Tested Every statistical claim about this manuscript rests on a transcription that two people hold and nobody else can inspect. The page images are free to everyone. We measured whether they are enough.
This is the only page on this site carrying a measurement of our own, and the measurement is a negative one. It is here because the alternative was to build a corpus we could not validate and publish statistics computed on it, which is precisely how work of this kind goes wrong.
I. There is no public transcription
Every statistical statement about a script depends on a transcription: a machine-readable record of which sign appears where. For the Rohonc Codex no such thing is publicly available. Each route that exists, checked:
| Route | State |
|---|---|
| The one open transcription project | Dead. The page returns 404. Its author's own note says he had got a long way and that life intervened, and it was never finished. |
| Delia Huegel's long-running study | Substantial interpretive work on the illustrations. Not a transcription dataset. |
| Király and Tokai | They necessarily hold a machine-readable transcription. The journal lists no supplementary data and it has not been published. |
| Page images | Available and public domain. 227 scans on the Internet Archive under CC0. Eight pages at higher resolution from Hamburg, 2015. |
So the images are free to anyone and the data is held by two people. This applies to claims in favour of the codex and against it equally, and it applies to the 2018 reading as much as to the three that failed. A result that cannot be independently checked is not being doubted when someone says so. It is being described.
II. What the public scans can and cannot support
The obvious response is to build the transcription from the images. We tested whether that is realistic.
The 227 public scans are two-page spreads, roughly 915 by 551 pixels, implying about 454 page images against a known 448. Working from local-mean binarisation, because a global threshold drowns in the ink showing through from the reverse of each leaf:
| Measurement | Result |
|---|---|
| Size of a glyph-sized connected component | median 12 × 14 pixels, 90th percentile 29 × 22 |
| Components per spread | median 443 |
| Rows per page, spreads left whole | median 6, against a published 9 to 14 |
| Rows per page, spreads split at the gutter | median 10, mean 10.2 |
| Pages landing inside the published 9 to 14 band | 52 per cent |
Splitting the spreads moved the row count into the published band, which says the approach is sound in principle. Fifty-two per cent per page is not sound enough to build a corpus on. Some of the low counts are certainly correct rather than mistaken, since roughly 87 pages carry illustrations and genuinely hold little text, and that has not been separated out.
Suppose the segmentation were perfect. It would give the location of every glyph and the identity of none. At twelve pixels across there is very little information per sign with which to separate several hundred classes, and, decisively, there is no ground truth against which a clustering could be checked. A transcription built this way could not be validated, and statistics computed on an unvalidated transcription would look exactly like a result while being an artefact of the clustering. So the corpus was not built.
III. What would change this
The transcription. Király and Tokai have one. Published, or shared with researchers, it would let their reading be tested independently by anyone, which is what their result deserves and has not had. It would also make every conventional structural measurement possible in an afternoon, and it would settle the sign inventory question that has been open since 1892.
Higher-resolution images. Eight pages exist at good resolution. At three to four times the linear resolution of the public scans, automatic sign clustering moves from marginal to genuinely promising, and the validation problem becomes tractable because a human can check a sample against legible images.
Until one of those arrives there is no Rohonc corpus and therefore no Rohonc result. The constraint is access to data, not analysis. That is an unglamorous conclusion and it is the honest one.
IV. What this site does not claim
- No sign in this manuscript is given a sound or a meaning anywhere on this site.
- Nothing here settles whether the codex is authentic or a fabrication. Both positions are reported and neither is endorsed.
- Nothing here confirms or refutes Király and Tokai. Their work has not been tested independently and cannot be until the transcription is available. That is a statement about access, not about quality.
- The measurements above are of the public scans, not of the manuscript. A better photograph would change them.