Answer in 30 seconds
Compare LLMs on the same documents, schema and scoring rules. Track field quality, failed runs, latency, provider cost and human correction time.
Write the decision before running models
Name the job and the pass condition first. You might need a pipeline for scanned supplier invoices that captures every line item with the correct amount. A general model leaderboard cannot answer that question for your documents.
Keep a controlled model test separate from an end-to-end pipeline test. A controlled comparison fixes the input, schema, prompt and scoring rules. A pipeline comparison lets each system use its normal preprocessing, but those differences must be reported. OCR text and page images are not equivalent inputs.
Keep the final document set out of prompt tuning
Adjust prompts and schemas on a development set. Make the final decision on documents you did not use for those changes. Include ordinary work and known difficult cases, and remove duplicates that make the test look broader than it is.
A pilot of 10 to 30 documents can expose obvious failures. It cannot support a universal accuracy claim. Record the suppliers, layouts, languages and scan conditions represented, then add documents where the result is still uncertain.
Version the source files and reviewed answers. Save the model identifier, provider, run date, prompt, schema, page selection, preprocessing and generation settings with each output. Keep the first prediction before a person corrects it.
Use one scorecard for every LLM
Score fields with the precision, recall and F1 method in the extraction guide. Record invalid output, missing pages, unsupported files, timeouts and retries separately. State whether you score the first attempt or the final pipeline output, and report the share of documents completed.
Run each model more than once. Report variation, median and slow-run latency, provider charges, hosted PE use and the minutes a reviewer spends fixing the result. Label cached runs separately because they measure cache behavior rather than fresh inference.
A useful pass rule could require every account identifier and invoice total to match, with no missing line items. Set the rule before seeing the results. When one model returns better data but needs more review time, show the trade-off instead of hiding it inside one average.
Run the same benchmark in Dynamite Docs
Choose models in AI Settings and run the fixed document set. The compatibility check confirms that a request works, but it does not measure extraction accuracy. The repository mock benchmark checks deterministic scoring behavior and says nothing about live provider quality. Export and retain the real outputs for comparison.
Method reference: OpenAI evaluation best practices. The worked examples and proposed test protocol above are illustrative guidance.
Test a document against ground truth