Method · ground truth

There is no ground truth for Afghan lithographs

Nobody has hand-transcribed a page of this collection, so there is nothing to score against. Here is how you manufacture a benchmark, and what it cannot tell you.

· Steps Ventures

Every claim about how well a machine reads a page rests on a page somebody already read by hand. For Afghan lithograph printing of the 1870s to the 1930s, that page does not exist. We looked hard. There is no page-level human transcription of the Afghanistan Digital Library anywhere, and as far as we can tell there is none for early Afghan lithograph print in general.

That is an awkward position. You have 69,624 pages of machine-read text and no way to say how good it is.

Manufacturing a reference

The way out is that some of these texts were reprinted in modern type. A modern edition is clean, and a model reads clean type very well. So you read the modern edition, align it page by page against the original scans, and have a model adjudicate every candidate pair with explicit instructions that shared author, topic and era are not evidence of a match. What survives is a set of verified image-to-text pairs, for the price of the compute.

For us that produced 102 aligned pairs of lithographed Kabul nastaliq, matching the Siraj al-akhbar newspaper against the 2008 reset edition of its own author’s collected articles. As far as we know that benchmark did not previously exist. The method transfers to any corpus with a modern reprint, and the code is open.

What the number actually means

On that benchmark, in the configuration we used to read the whole collection, the median character error rate is 0.361. Roughly two thirds of characters agree. That sounds bad and it overstates how bad the reading is, for a reason that has nothing to do with the machine.

A modern editor normalises spelling, adds punctuation, and regularises archaic forms. A Kabul printing may follow a different recension of the text entirely. So a chunk of that 0.361 is the distance between a 1911 lithograph and a 2008 reprint, not a misreading. We measured that effect separately at about 0.19, and here is the part it is tempting to skip: we measured it on a different pairing, classical verse against a canonical text, not journalism against its author’s reprint. Carrying it across is an analogy. It is not a measurement. Subtracting it entirely would imply about 83% of characters correct, and nobody has verified that.

Why we publish the uncomfortable version

It would be easy to lead with a better number. We have one. On human-transcribed Persian nastaliq manuscripts the same model family reads at about 93% of characters. That figure is real and it is about a different corpus, in a different medium, and quoting it as this archive’s accuracy would be a lie of exactly the kind that is hardest to catch, because every individual word of it is true.

The first person to check will be a subject specialist with the page open beside them. If the published number does not survive their first ten minutes, nothing else on the site is worth anything either.

The ask

The single most valuable thing anyone could give this project is a handful of hand-transcribed pages of Afghan lithograph print. Not a thousand. Twenty would change what we can say. If your institution holds any, we would gladly hand over the full gold set and the code that produced it in exchange.

Until then the honest description of the text is a finding aid rather than an edition. It is good enough to find the page you want across 69,624 of them. It is not good enough to quote. Read the image.