71%
of lines contain at least one model mistake (114 and 111 of 161 lines)
Human-verified eval dataset · Historical Urdu Nastaliq
We ran a blind test of Claude Opus 5.5 and Claude Fable 5.1 on a 1908 lithographed Urdu dictionary, scored word by word against ground truth typed by a native expert. Every disagreement was ruled against the scan. The sample, the harness method and the word-level results are published here.
71%
of lines contain at least one model mistake (114 and 111 of 161 lines)
55 · 50
words each model got wrong with no doubt flagged (Opus · Fable)
1 in 26
characters wrong at best. Ground-truth grade (99.995%) is 1 in 20,000.
130 · 113
diacritics written by the katib that each model dropped
What the test shows
Both models were told to flag any word they were unsure of. They still misread 55 and 50 words without a flag. Nothing in the output marks them as wrong.
The katib wrote pre-Partition Urdu in full rasm-ul-khat, with shadd, jazm, zer and pesh. The models dropped 130 and 113 of these marks, rewriting 1908 as modern print.
Urdu has 246 million speakers (Ethnologue 2025), yet it appears in none of the 52 disclosed AI data-licensing deals compiled by Neudata. Verified ground truth for historical Nastaliq is scarce, so there is little to train on and nothing standard to measure against.
Results
Reading accuracy counts misread, missing and invented words. Silent errors are reading errors with no doubt flag.
| Opus 5.5 reading accuracy | Misread | Silent | Fable 5.1 reading accuracy | Misread | Silent | |
|---|---|---|---|---|---|---|
| Page 12 | 98.71% | 6 | 4 | 97.20% | 13 | 6 |
| Page 13 | 96.15% | 37 | 22 | 95.00% | 48 | 26 |
| Page 14 | 95.70% | 41 | 29 | 96.12% | 37 | 18 |
| All 3 pages | 96.47% | 84 | 55 | 95.88% | 98 | 50 |
The eval harness
Provenance
Farhang-i-Asifiyah, vol. 4 (1908), compiled by Syed Ahmad Dehlvi (1846–1918). Public domain in the US, Pakistan and India.
Scan: archive.org Farhang-i-afiyah_1908_amaduoft, pages 12–14.
The ground truth is typed by people from the scan. Model outputs appear only as scored evaluation evidence.
124 titles sourced (49,421 pages) in Urdu, Punjabi, Sindhi, Pashto, Balochi and Saraiki, selected as public-domain works published before 1931.
Data format
A real record
{
"page": 13,
"model": "claude-fable-5-1",
"band": "R3",
"truth": "مدھ",
"model_output": "مدہ",
"result": "wrong_word",
"model_flagged_doubt": false,
"reading_error": true,
"scan_crop": "p13_R3.png"
}
Work with us
Pilot: a 250-page held-out set of historical Nastaliq, scored against your model with this harness. We send the scans, you send your model's output, and we return the scorecard and full error log. Your model never sees the truth.