Skip to content
PAIR
EN
Request a demo Demo

What an AI model can read on a paper form

What is measured today about handwriting recognition, why a low error per character is not a low error per field, and what is not in the ink.

Hands in work gloves holding a phone that photographs a handwritten paper form, in the field.
AI-generated image from our own references.

1.75 %

error per character on modern handwriting (Crosilla, Klic and Colavizza, 2025). Per field it is not 1.75 %.

The question comes up in almost every conversation, and it arrives already decided: if an artificial intelligence model can read anything today, why not feed it the scanned sheets and skip the digital field work? It is a good question and it deserves an answer with numbers, not enthusiasm. We went looking for the numbers.

What is measured

The most useful work we found is by Crosilla, Klic and Colavizza, published in 2025 in the Journal of Documentation under the title Benchmarking large language models for handwritten text recognition. It compares multimodal language models against Transkribus, the trade’s specialised engine, on handwritten corpora in English, French, German and Italian. The open version is at arXiv:2503.15195.

On modern English handwriting —the IAM corpus— the general-purpose models win comfortably: 1.71 % character error rate for GPT-4o-mini and 1.75 % for Claude 3.5 Sonnet, against 9.13 % for Transkribus’s specialised model. On historical Italian handwriting the sign flips but not the order of magnitude: the best model lands at 20.55 % character error rate.

Two warnings before drawing a conclusion. The first is that the study measures language: performance drops outside English, and none of its corpora is Spanish. The second is that none of those corpora is an inspection form, with checkboxes, site acronyms, trade abbreviations and an «N/A» written six different ways.

Why 1.75 % is not 1.75 %

Character error is a character metric, and a form is read by field. If errors fell one by one and at random —they do not; they cluster— a six-character field read at 1.75 % character error would come out wrong slightly more than one time in ten. The sum is ours, it is the worst case and it serves one purpose only: to show that a low rate per character is not a low rate per field, which is the unit that later goes into a report, a statistic or a lawsuit.

A handwritten inspection sheet, stained with dust, and beside it a phone with the same sheet photographed, on a field table.

AI-generated image from our own references.

What the model cannot do about itself

The same study tests something often proposed as a solution: asking the model to review its own transcription. The authors’ conclusion is that self-correction produces no substantial improvement over the initial prediction; the gains are inconsistent, and in the open models the second pass sometimes increased the error rate.

That has a practical consequence. «Let the model check itself» is not quality control: it is the same reading again. The control has to come from outside —from a human, from a rule, from a second piece of data that does not come out of the same ink.

What is not in the ink

There is a limit that does not depend on the model and that no future version will lift, because it is not a reading problem.

A scanned sheet says what was written. It does not say who wrote it, or when, or where that person was, or against which version of the form they answered, or whether the photograph attached three days later belongs to that finding or to the one next to it. None of that data is on the paper, so no model can recover it: it would infer it, which is something else, and in a record that may end up in an inspection that difference is the one that matters.

That is why the short answer to the opening question is that transcription does not create provenance. Capturing the structured data at the moment and in the place it happens is cheaper than rebuilding it afterwards, and above all it is the only thing that leaves a trace of who and when.

Where it does help

None of the above says the model is useless. It says where to put it. On data already captured and already traced, a model does well things a person does slowly: drafting a description, grouping similar findings from different months, proposing a classification somebody then confirms, finding the inspection from two years ago that resembles today’s. In all of them there is a person who decides, and a prior record to verify against.

What we did not measure

We have not built our own benchmark on Spanish-language site forms, and the figures above are not from that material: they are from academic corpora, in other languages and in other handwriting. They serve to bound an order of magnitude and to dismiss the idea that the problem is already solved. They do not serve to promise a percentage on a Chilean sheet written with gloves on.

Sources

Crosilla G., Klic L. and Colavizza G., Benchmarking large language models for handwritten text recognition, Journal of Documentation, vol. 81 no. 7 (2025), pp. 334-354: doi 10.1108/JD-03-2025-0082 · open version at arXiv:2503.15195.

All the news