Answer in 30 seconds
Measure OCR accuracy against a reviewed transcript with character error rate and word error rate, then inspect the mistakes that can damage extracted data.
Build ground truth before you test OCR
Choose pages that look like the work you actually receive. Keep digital PDFs, scans and phone photos in separate groups so clean files cannot hide failures on faint receipts or rotated statements. Record the file, page number and reading order for every reference transcript.
Have one person transcribe each page and another check ambiguous text against the image. Mark unreadable regions instead of guessing. Write down how the test treats whitespace, punctuation, capitalization and Unicode, then use those rules for every OCR engine.
Calculate character and word error rates
Character error rate, CER, equals substitutions plus deletions plus insertions, divided by the number of reference characters. Word error rate, WER, applies the same edit-distance calculation to words. Lower is better. Extra inserted text can push either rate above 100%.
Worked example: the reference is "Total 125.00" and the prediction is "Total 128.00". Counting the space, there are 12 reference characters and one substitution, so CER is 1/12 = 8.33%. With whitespace tokenization there are two words and one wrong word, so WER is 1/2 = 50%. This is an illustrative calculation, not a Dynamite Docs benchmark result.
For a corpus score, add all edit counts and reference lengths before dividing. Keep the page-level results too. If a reference is empty, the denominator is zero, so report inserted text separately and state the scorer convention rather than recording a perfect score.
Find the OCR errors that change the data
A wrong digit in a total can matter more than dozens of harmless punctuation errors. Keep a separate error register for account numbers, dates, amounts and identifiers. Do not strip decimal points or minus signs during normalization simply to improve the score.
CER and WER cannot tell whether a value landed in the right field or table row. A transcript may contain every word and still mix two columns. Use the extraction accuracy guide to score field assignments, then use the LLM benchmark guide for complete pipeline comparisons.
Method reference: Hugging Face CER metric definition and limitations. The worked examples and proposed test protocol above are illustrative guidance.
Test a document against ground truth