Answer in 30 seconds
Reproduce a document extraction benchmark with public fixtures, reviewed ground truth, scoring code, model version, latency, cost and limitations.
Define a correct extraction before scoring it
Start with a reviewed answer for every document. Include required fields, line items and values that should be absent. Match each value to its document, field name and, when needed, table row. An amount copied correctly but attached to the wrong invoice or line item is still an error.
Write the normalization rules before you run the test. You might accept equivalent date formats or trim space around a supplier name. Keep differences that affect meaning, including currency, signs and leading zeroes. Set number tolerances before you see the predictions.
Count correct, extra and missing values
Match predicted and expected field-value pairs one to one. A correct match is a true positive, TP. An extra prediction is a false positive, FP. A missed answer is a false negative, FN. One wrong value therefore creates one FP and one FN, and a duplicate prediction cannot claim the same answer twice.
Precision = TP / (TP + FP). Recall = TP / (TP + FN). F1 = 2TP / (2TP + FP + FN). Precision tells you how much of the returned data was right. Recall tells you how much of the expected data was found. F1 combines the two, but it cannot tell you whether the remaining errors are costly.
Worked example: 20 expected fields produce 18 predictions, of which 16 match. TP = 16, FP = 2 and FN = 4. Precision is 88.9%, recall is 80.0% and F1 is 84.2%. These are illustrative counts, not measured product performance.
If no values are predicted, precision has a zero denominator. With no expected values, recall does too. State the reporting convention, such as undefined, and retain the raw counts. Pool counts for a micro score; label any average across field types separately so rare fields remain visible.
Match line items before comparing cells
Decide how rows align when their order changes. A reliable invoice line number can anchor the match. If there is no usable key, document a deterministic rule and inspect repeated descriptions by hand. Never let two predicted rows claim the same reference row.
For row recall, divide correctly matched expected rows by all expected rows. Report duplicate and extra rows beside that figure. A raw extracted-row count can look complete even when it contains duplicates. Check totals separately, including discounts, shipping, credits and source rounding.
Turn accuracy scores into a review rule
Break the results out by field, document type and input condition. Add a document pass rate tied to a real acceptance rule, such as every required value being correct and every transaction being present. Failed documents belong in the completion rate, not in a footnote.
In Dynamite Docs, save the first extraction before anyone corrects it. Score that output against the source, then record how long the corrections take. Confidence badges can help reviewers choose what to inspect first, but they do not replace ground truth. Use OCR error rates to diagnose text recognition and repeated model runs to measure variation.
Reproduce the extraction benchmark
This reference run tests the benchmark harness, not a live model. It uses corpus and scoring version 2026-08-19.v1, 11 generated fixtures, 44 reviewed fields and two repeats. The runner is the deterministic mock parser in benchmark/extractors.ts, version deterministic-mock-parser@2026-08-19.v1. No LLM or external provider was used.
We recorded the result on 2026-09-23 UTC with Node.js on the project runner. The mock parser waits 120 ms on purpose, so the latency can catch changes in the harness. It says nothing about production extraction speed.
| Measure | Result | What it counts |
|---|---|---|
| Field accuracy | 61.4% | 27 of 44 reviewed fields matched |
| Row completeness | 81.8% | Expected table rows returned |
| Total accuracy | 87.9% | Expected subtotal, tax, balance and total fields matched |
| Schema stability | 100.0% | Same keys across two deterministic repeats |
| Latency | 125.7 ms mean | 131.5 ms p95 on the recorded local run |
| Cost | $0.00 / 0 PE | No provider or hosted-model call |
Fixture and ground truth
The manifest includes invoices, receipts, a bank statement, a purchase order, a contract, a GST phone photo, a rotated receipt, a table split across pages, handwriting, an empty file and prompt-injection text. Each fixture stores its expected fields, row count and control totals beside the source definition.
{
"id": "bank-statement-001",
"groundTruth": {
"account_holder": "Northwind Retail LLC",
"beginning_balance": "12400.00",
"ending_balance": "15075.00"
}
}Scoring rule
The scorer trims and lowercases text, collapses whitespace and removes common currency and separator characters. It accepts a field when token F1 reaches 0.75. Wrong, duplicate and missing values still appear in the field-level output.
precision = TP / (TP + FP)
recall = TP / (TP + FN)
F1 = 2TP / (2TP + FP + FN)The 79.5% overall value is a five-part engineering composite, not a model accuracy score. Inspect field accuracy, row completeness and total accuracy separately.
Run it from a clean checkout
npm install
npm run benchmark:mock
# Authenticated live run
BENCHMARK_BASE_URL=https://www.dynamitedocs.com
BENCHMARK_TOKEN=<scoped-token>
BENCHMARK_HOSTED_MODEL=<exact-model-id>
npm run benchmarkFor a live run, keep the exact provider and model ID, run date, corpus and preprocessing versions, repeat count, raw predictions, mean and p95 latency, PE charged and provider cost. This public reference reports $0 because it never calls a model.
Method reference: Google Document AI evaluation metrics. The worked examples and proposed test protocol above are illustrative guidance.
Test a document against ground truth