Audit Document Extraction Without Rekeying Evidence

Dynamite Docs, 2026-08-12

What audit document extraction should prepare

Audit document extraction prepares client files for testing without making the audit conclusion. It turns trial balances, general ledger reports, bank confirmations, debt schedules, invoices, and other support into structured rows. The original document remains the source. The table makes its fields easier to sort, recalculate, compare, and review.

A usable row needs more than an amount. It should identify the entity, period, document type, source filename, page, account or business key, currency, extracted value, review state, reviewer, and note. The exact fields come from the procedure. A confirmation test needs different identifiers from a trial balance tie-out.

Keep extraction, validation, and audit judgment separate. Software can organize fields and expose arithmetic differences. The engagement team decides whether the evidence is relevant and reliable, whether more work is needed, and what conclusion the procedure supports.

  • Input: original client files plus a complete request or population list.
  • Preparation: fields extracted under a procedure-specific schema.
  • Control: counts, totals, dates, duplicates, and missing items checked.
  • Review: exceptions resolved by an assigned person with a recorded note.
  • Output: accepted rows and an exception log linked to stable source identifiers.

Start with the procedure, not the PDF layout

Define the question the working paper must answer before choosing columns. For a trial balance tie-out, capture account code, account description, debit, credit, net balance, reporting period, entity, and source page. For bank confirmations, capture the responding institution, account identifier, balance, currency, confirmation date, restrictions, and response status. A loan schedule may need principal, interest rate, maturity date, payment date, and covenant fields.

Do not reproduce every printed label just because it appears on the page. Capture the fields needed for the test and enough source context to explain them. This keeps the table focused while preserving the original file for details outside the schema.

Use stable identifiers that survive sorting and export. File names alone can change, so pair the file with an engagement identifier, request-list item, source page, and row or account key. When one document supplies several tested values, each row should still point to the same source record.

  • Trial balance: entity, period, account code, description, debit, credit, and net balance.
  • Confirmation: counterparty, account identifier, date, currency, balance, restriction, and response state.
  • Schedule: agreement or facility, period, opening value, movement, closing value, and relevant dates.
  • Sample support: population key, sample number, document identifier, tested amount, and exception state.

Prove the population is complete before testing rows

A perfectly copied sample does not fix an incomplete population. Reconcile the files received with the request list before relying on the extracted table. Compare document counts, page counts, entity and period coverage, sequence numbers where they exist, and any control total supplied by the client.

Detect duplicates using more than the filename. A renamed copy can still repeat the same file, while two legitimate documents can share a generic name such as statement.pdf. Combine a file hash with business keys such as account, date, amount, invoice number, or confirmation reference, then send possible duplicates to review.

For numeric populations, total the extracted rows and compare that result with the source report or agreed control figure before selecting samples. Investigate differences rather than adjusting the spreadsheet to force agreement. Common causes include omitted pages, repeated headers treated as data, sign errors, subtotal rows mixed into detail, and rows split across page breaks.

Extract values while preserving source context

Process digital PDFs with their text layer intact when possible. Scans, image-only PDFs, and photographed documents require audit document OCR or a visual model, which adds more uncertainty around small type, signs, decimal points, and handwritten changes. Keep the original bytes regardless of the extraction method.

Store the extracted value separately from any corrected value. Record the reason, reviewer, and time for each correction. If a balance changes from 15,800 to 15,300 after review, the working file should show both values and the source used to resolve the difference. Replacing the first value in place erases useful evidence about the preparation step.

A page reference is the minimum useful pointer for a multi-page file. Add a table name, row label, account code, or visible section when the page contains many similar values. If the review interface highlights a source region, treat that as a review aid. Keep durable text identifiers in the exported working paper because screen coordinates may not travel with the file.

Use recalculation and tie-outs to find extraction errors

Validation should test the relationships already present in the source. For a trial balance, check whether debit and credit columns total under the report's convention and whether account balances agree with the control figure used by the procedure. For roll-forward schedules, test opening balance plus movements against closing balance. For invoices or statements used as support, recalculate printed subtotals where the source supplies the needed components.

Do not apply one generic formula to every document. Reports use different sign conventions, netting rules, and summary layouts. Record the rule used for each test and preserve the printed control total. A failed check may point to an extraction error, a client-prepared report issue, or a misunderstanding of the report. The reviewer decides which.

Scan for structural failures as well as arithmetic ones. Check the first and last row on every page, repeated headers, wrapped descriptions, blank account codes, unexpected periods, mixed currencies, duplicate keys, and missing sequence values. A table can total correctly and still contain a duplicated row paired with an omitted row of the same amount.

Route exceptions without hiding unresolved evidence

Give each row a clear state such as Extracted, Needs review, Corrected, Accepted, Rejected, or Missing evidence. Add narrower reason codes when they help the team act, such as Unreadable value, Control total mismatch, Possible duplicate, Wrong period, Missing page, or Source not received.

Keep unresolved rows in an exception log. Do not delete them from the population just because they cannot enter the final schedule. The log should show the source item, issue, owner, requested action, latest status, and resolution. This lets reviewers distinguish incomplete work from accepted evidence.

Approval should record who reviewed the row and what they checked. A confidence score can help prioritize attention, but it does not approve a value. High-confidence fields can still be wrong, especially when the source itself contains an error or the schema assigned the right number to the wrong label.

  • Unreadable source: request a clearer file or leave the value unresolved.
  • Population gap: return to the request list and identify the missing document or period.
  • Control mismatch: isolate omitted, repeated, or mis-signed rows before sampling.
  • Ambiguous correction: retain both values and escalate the judgment to the assigned reviewer.

Keep the working paper reviewable and access controlled

Treat client files according to the engagement's confidentiality, retention, and access rules. Limit the workspace and exports to the people who need them. Avoid placing full account numbers, tax identifiers, or personal data in filenames, chat messages, or reviewer notes when a masked reference will do.

Provider choice matters when an extraction step sends document content to an AI service. Confirm the provider, region, retention terms, and training policy that apply to the selected model. Dynamite Docs also supports user-supplied provider keys and a local Ollama companion. Local processing can keep the model step on the user's machine, but teams still need to control the original files, exports, backups, and access to that machine.

Keep versions of the source list, extraction output, corrections, accepted working paper, and exception log. Follow the firm's approved retention policy and the professional standards that govern the engagement. This article explains a preparation method, not a retention period or audit requirement for every jurisdiction.

What audit standards say about evidence and documentation

PCAOB AS 1105 says auditors must obtain sufficient appropriate audit evidence. It treats sufficiency as quantity and appropriateness as relevance and reliability. The standard also requires auditors to test the accuracy and completeness of information produced by the company, or test the controls over that information, when they use it as audit evidence.

PCAOB AS 1215 requires audit documentation to identify the procedures performed, evidence obtained, conclusions reached, preparer, and reviewer. For procedures involving inspected documents, it also calls for identification of the items inspected. These requirements explain why an extracted number without a stable source identifier is a weak working-paper record.

The applicable framework depends on the engagement and jurisdiction. Use the standards and firm methodology that govern the actual work. Extraction software prepares information for those procedures. It does not decide sufficiency, reliability, sample design, or the audit conclusion.

  • PCAOB AS 1105, Audit Evidence (Source: PCAOB AS 1105, Audit Evidence)
  • PCAOB AS 1215, Audit Documentation (Source: PCAOB AS 1215, Audit Documentation)

Frequently asked questions about audit document extraction

These answers define where working paper data extraction ends and professional review begins.

  • Does extracted data replace the original document? No. Keep the source as the evidence and treat the structured row as a reviewable representation prepared for a procedure.
  • What should happen when a value is unreadable? Leave it unresolved, record the source and issue, and request review or better evidence. Do not guess.
  • Can extraction software perform the audit conclusion? No. It can organize and check data. The engagement team owns procedure design, evidence evaluation, judgment, and conclusions.
  • Should every extracted field be reviewed? Set review rules through the engagement methodology and risk assessment. At minimum, investigate failed controls, uncertain fields, structural exceptions, and the items selected for the procedure.
  • Can a confidence score prove that a value is correct? No. It is a routing signal, not audit evidence or approval. Test important values against the source and the procedure's control checks.

Prepare one traceable working paper before scaling the batch

Start with one representative source set and the procedure it supports. Define the schema, reconcile the population, extract the rows, preserve source identifiers, run control checks, and resolve exceptions. Export the accepted rows only after a reviewer can follow a value back to the document and understand every correction.

Dynamite Docs can organize varied documents into reviewable tables and export accepted data to XLSX, CSV, JSON, or Google Sheets. Test the workflow on the most difficult file in the batch, not only the cleanest PDF. If the working paper stays complete and traceable through that review, reuse the schema for the remaining files.

Sources checked for this note

  1. PCAOB AS 1105, Audit Evidence
  2. PCAOB AS 1215, Audit Documentation

Related workflow: Accounting document automation and data extraction.

Keep reading

Try it yourself. Upload a PDF, scan, or image and let Dynamite Docs infer the schema.

Loading Dynamite Docs… This page is taking longer than expected. Reload page.