Precision / recall / F1

Stress-testing multi-page table extraction on complex real-world layouts

Reproduce a document extraction benchmark with public fixtures, reviewed ground truth, scoring code, model version, latency, cost and limitations.

Answer in 30 seconds

Reproduce a document extraction benchmark with public fixtures, reviewed ground truth, scoring code, model version, latency, cost and limitations.

01

Define a correct extraction before scoring it

Start with a reviewed answer for every document. Include required fields, line items and values that should be absent. Match each value to its document, field name and, when needed, table row. An amount copied correctly but attached to the wrong invoice or line item is still an error.

Write the normalization rules before you run the test. You might accept equivalent date formats or trim space around a supplier name. Keep differences that affect meaning, including currency, signs and leading zeroes. Set number tolerances before you see the predictions.

02

Count correct, extra and missing values

Match predicted and expected field-value pairs one to one. A correct match is a true positive, TP. An extra prediction is a false positive, FP. A missed answer is a false negative, FN. One wrong value therefore creates one FP and one FN, and a duplicate prediction cannot claim the same answer twice.

Precision = TP / (TP + FP). Recall = TP / (TP + FN). F1 = 2TP / (2TP + FP + FN). Precision tells you how much of the returned data was right. Recall tells you how much of the expected data was found. F1 combines the two, but it cannot tell you whether the remaining errors are costly.

Worked example: 20 expected fields produce 18 predictions, of which 16 match. TP = 16, FP = 2 and FN = 4. Precision is 88.9%, recall is 80.0% and F1 is 84.2%. These are illustrative counts, not measured product performance.

If no values are predicted, precision has a zero denominator. With no expected values, recall does too. State the reporting convention, such as undefined, and retain the raw counts. Pool counts for a micro score; label any average across field types separately so rare fields remain visible.

03

Match line items before comparing cells

Decide how rows align when their order changes. A reliable invoice line number can anchor the match. If there is no usable key, document a deterministic rule and inspect repeated descriptions by hand. Never let two predicted rows claim the same reference row.

For row recall, divide correctly matched expected rows by all expected rows. Report duplicate and extra rows beside that figure. A raw extracted-row count can look complete even when it contains duplicates. Check totals separately, including discounts, shipping, credits and source rounding.

04

Turn accuracy scores into a review rule

Break the results out by field, document type and input condition. Add a document pass rate tied to a real acceptance rule, such as every required value being correct and every transaction being present. Failed documents belong in the completion rate, not in a footnote.

In Dynamite Docs, save the first extraction before anyone corrects it. Score that output against the source, then record how long the corrections take. Confidence badges can help reviewers choose what to inspect first, but they do not replace ground truth. Use OCR error rates to diagnose text recognition and repeated model runs to measure variation.

05

Reproduce the extraction benchmark

This reference run tests the benchmark harness, not a live model. It uses corpus and scoring version 2026-08-19.v1, 11 generated fixtures, 44 reviewed fields and two repeats. The runner is the deterministic mock parser in benchmark/extractors.ts, version deterministic-mock-parser@2026-08-19.v1. No LLM or external provider was used.

We recorded the result on 2026-09-23 UTC with Node.js on the project runner. The mock parser waits 120 ms on purpose, so the latency can catch changes in the harness. It says nothing about production extraction speed.

Deterministic reference result, 11 fixtures and 2 repeats
MeasureResultWhat it counts
Field accuracy61.4%27 of 44 reviewed fields matched
Row completeness81.8%Expected table rows returned
Total accuracy87.9%Expected subtotal, tax, balance and total fields matched
Schema stability100.0%Same keys across two deterministic repeats
Latency125.7 ms mean131.5 ms p95 on the recorded local run
Cost$0.00 / 0 PENo provider or hosted-model call

Fixture and ground truth

The manifest includes invoices, receipts, a bank statement, a purchase order, a contract, a GST phone photo, a rotated receipt, a table split across pages, handwriting, an empty file and prompt-injection text. Each fixture stores its expected fields, row count and control totals beside the source definition.

{
  "id": "bank-statement-001",
  "groundTruth": {
    "account_holder": "Northwind Retail LLC",
    "beginning_balance": "12400.00",
    "ending_balance": "15075.00"
  }
}

Scoring rule

The scorer trims and lowercases text, collapses whitespace and removes common currency and separator characters. It accepts a field when token F1 reaches 0.75. Wrong, duplicate and missing values still appear in the field-level output.

precision = TP / (TP + FP)
recall    = TP / (TP + FN)
F1        = 2TP / (2TP + FP + FN)

The 79.5% overall value is a five-part engineering composite, not a model accuracy score. Inspect field accuracy, row completeness and total accuracy separately.

Run it from a clean checkout

npm install
npm run benchmark:mock

# Authenticated live run
BENCHMARK_BASE_URL=https://www.dynamitedocs.com
BENCHMARK_TOKEN=<scoped-token>
BENCHMARK_HOSTED_MODEL=<exact-model-id>
npm run benchmark

For a live run, keep the exact provider and model ID, run date, corpus and preprocessing versions, repeat count, raw predictions, mean and p95 latency, PE charged and provider cost. This public reference reports $0 because it never calls a model.

Stress-testing multi-page table extraction on complex real-world layouts: reviewable document extractionReview the source document and extracted rows in one workspace.
Stress-testing multi-page table extraction on complex real-world layouts: table and line item extractionCheck quantities, descriptions, and prices row by row.
Stress-testing multi-page table extraction on complex real-world layouts: human-in-the-loop validationResolve flagged values before approving the extraction.
Stress-testing multi-page table extraction on complex real-world layouts: Excel, CSV and JSON exportDownload reviewed data as CSV, Excel, or JSON.

Method reference: Google Document AI evaluation metrics. The worked examples and proposed test protocol above are illustrative guidance.

Test a document against ground truth

Choose what to measure

Text recognition, model behavior and structured field quality answer different questions. Use the same reviewed documents to connect the results.

All accuracy testing methods

Apply the test in Dynamite Docs

Choose a document workflow, run your sample, and compare the result with your reviewed answer.

  • Invoices

    Check identifiers, tax, totals and line items.

  • Receipts

    Test faint text, tips, dates and merchant names.

  • Bank statements

    Check transaction coverage, signs and balances.

  • PDFs

    Separate text-layer, scanned and mixed-page results.

  • API

    Retain outputs and track failures across repeatable runs.

  • PDF tables

    Check table boundaries, page breaks, row counts and totals.

  • Statement to Excel

    Test transaction coverage, signs and balance reconciliation.

  • BUSY invoices

    Verify GSTIN, HSN or SAC, tax splits and bill totals.

Loading Dynamite Docs… This page is taking longer than expected. Reload page.