Model comparison / repeated runs

Field extraction precision and recall metrics across frontier LLMs

Compare LLMs on the same documents, schema and scoring rules. Track field quality, failed runs, latency, provider cost and human correction time.

Answer in 30 seconds

Compare LLMs on the same documents, schema and scoring rules. Track field quality, failed runs, latency, provider cost and human correction time.

01

Write the decision before running models

Name the job and the pass condition first. You might need a pipeline for scanned supplier invoices that captures every line item with the correct amount. A general model leaderboard cannot answer that question for your documents.

Keep a controlled model test separate from an end-to-end pipeline test. A controlled comparison fixes the input, schema, prompt and scoring rules. A pipeline comparison lets each system use its normal preprocessing, but those differences must be reported. OCR text and page images are not equivalent inputs.

02

Keep the final document set out of prompt tuning

Adjust prompts and schemas on a development set. Make the final decision on documents you did not use for those changes. Include ordinary work and known difficult cases, and remove duplicates that make the test look broader than it is.

A pilot of 10 to 30 documents can expose obvious failures. It cannot support a universal accuracy claim. Record the suppliers, layouts, languages and scan conditions represented, then add documents where the result is still uncertain.

Version the source files and reviewed answers. Save the model identifier, provider, run date, prompt, schema, page selection, preprocessing and generation settings with each output. Keep the first prediction before a person corrects it.

03

Use one scorecard for every LLM

Score fields with the precision, recall and F1 method in the extraction guide. Record invalid output, missing pages, unsupported files, timeouts and retries separately. State whether you score the first attempt or the final pipeline output, and report the share of documents completed.

Run each model more than once. Report variation, median and slow-run latency, provider charges, hosted PE use and the minutes a reviewer spends fixing the result. Label cached runs separately because they measure cache behavior rather than fresh inference.

A useful pass rule could require every account identifier and invoice total to match, with no missing line items. Set the rule before seeing the results. When one model returns better data but needs more review time, show the trade-off instead of hiding it inside one average.

04

Run the same benchmark in Dynamite Docs

Choose models in AI Settings and run the fixed document set. The compatibility check confirms that a request works, but it does not measure extraction accuracy. The repository mock benchmark checks deterministic scoring behavior and says nothing about live provider quality. Export and retain the real outputs for comparison.

Field extraction precision and recall metrics across frontier LLMs: reviewable document extractionReview the source document and extracted rows in one workspace.
Field extraction precision and recall metrics across frontier LLMs: table and line item extractionCheck quantities, descriptions, and prices row by row.
Field extraction precision and recall metrics across frontier LLMs: human-in-the-loop validationResolve flagged values before approving the extraction.
Field extraction precision and recall metrics across frontier LLMs: Excel, CSV and JSON exportDownload reviewed data as CSV, Excel, or JSON.

Method reference: OpenAI evaluation best practices. The worked examples and proposed test protocol above are illustrative guidance.

Test a document against ground truth

Choose what to measure

Text recognition, model behavior and structured field quality answer different questions. Use the same reviewed documents to connect the results.

All accuracy testing methods

Apply the test in Dynamite Docs

Choose a document workflow, run your sample, and compare the result with your reviewed answer.

  • Invoices

    Check identifiers, tax, totals and line items.

  • Receipts

    Test faint text, tips, dates and merchant names.

  • Bank statements

    Check transaction coverage, signs and balances.

  • PDFs

    Separate text-layer, scanned and mixed-page results.

  • API

    Retain outputs and track failures across repeatable runs.

  • PDF tables

    Check table boundaries, page breaks, row counts and totals.

  • Statement to Excel

    Test transaction coverage, signs and balance reconciliation.

  • BUSY invoices

    Verify GSTIN, HSN or SAC, tax splits and bill totals.

Loading Dynamite Docs… This page is taking longer than expected. Reload page.