Document extraction API for PDFs, invoices and OCR

IN + OUT: /api/v1/extractions, /api/v1/infer, /api/tokens

Content reviewed .

Put the document extraction API inside your product, CI job or internal tool. Create a scoped key, send it as Authorization: Bearer dd_live_<key> or X-API-Key, and POST the file to /api/v1/infer. The versioned API returns structured rows without a browser session, so your code can validate the response and decide what happens next.

Open this integration in the app or read the public API and webhook guide.

A manual export button is not an API

  • If a person must click Export, every downstream job waits for that person.
  • Scripts need a stable credential with a narrow scope. A browser session cookie is the wrong tool.
  • The response needs predictable rows, metadata and errors that code can validate.
  • A token that can do everything becomes a long-lived liability once it lands in a codebase.

How the API works

01. Create a scoped key

From the app header under Webhooks & API (or /app?modal=webhooks), mint a scoped key. Keys look like dd_live_<64 hex>. The server stores only their SHA-256 hash, and the full secret is shown exactly once. Manage (list, revoke) them from the same panel.

02. Pull extractions from /api/v1

GET /api/v1/extractions lists your tables with pagination and filtering; GET /api/v1/extractions/:id/export?format=csv|json|xlsx returns one in the format your pipeline consumes. Send your key via the Authorization: Bearer <key> header or X-API-Key header.

03. Or infer from a document

POST /api/v1/infer uploads a document (multipart/form-data) and runs the same extraction engine as the app, returning the structured table in the response. It counts as one API call. Successful hosted pages and Free or Starter BYOK pages draw from the monthly PE budget; Pro or Scale BYOK pages cost 0 PE. Extractions persist and trigger webhooks. The interactive POST /api/extract/infer remains for session-authenticated calls.

04. Feed it into your systems

The JSON shape drops straight into your pipeline: a procurement system ingesting invoice rows, a finance tool importing bank statement transactions. The shape is consistent, so one consumer handles every document type.

What the API exposes

The versioned API exposes the calls a document pipeline needs:

  • Scoped keys: hashed at rest, revocable, secret shown once
  • Authentication: Authorization: Bearer or X-API-Key
  • Versioned endpoints: GET /api/v1/extractions, GET :id/export, POST /api/v1/infer
  • Pagination & filters: page, limit, sort, order, doc_type
  • Formats: CSV (UTF-8 BOM), JSON, XLSX
  • Inference: POST /api/v1/infer (multipart upload)
  • Usage metering: GET /api/v1/usage (calls vs monthly limit)
  • Rate limiting: 60/min per key, 20/min inference
A sample invoice open in the Dynamite Docs source viewer with extraction controls
The source document stays visible while you configure extraction.

A realistic example

A two-call integration, start to finish:

CallPurposeReturns
POST /api/tokens (session auth)Create a scoped API keydd_live_<64 hex>, shown once
GET /api/v1/extractions/<id>/export?format=json (Bearer / X-API-Key)Pull an extraction with the key{ columns, rows, metadata, docType }

Request and response examples

Export extraction as JSON

GET /api/v1/extractions/:id/export?format=json

curl -X GET \
  -H "Authorization: Bearer dd_live_a8f92b7c4d1e0394..." \
  "https://app.dynamitedocs.com/api/v1/extractions/ext_01j5x9b3/export?format=json"
{
  "id": "ext_01j5x9b3",
  "fileId": "file_8829f0",
  "fileName": "vendor_invoice_1092.pdf",
  "docType": "Invoice",
  "label": "August Operations",
  "columns": [
    "Invoice Number",
    "Date",
    "Vendor",
    "Description",
    "Line Total",
    "Tax",
    "Amount Due"
  ],
  "rows": [
    [
      "INV-1092",
      "2026-08-01",
      "Acme Industrial Ltd",
      "Precision CNC milling components",
      "$1,250.00",
      "$100.00",
      "$1,350.00"
    ]
  ],
  "metadata": {
    "vendor": "Acme Industrial Ltd",
    "invoice_number": "INV-1092",
    "total": "1350.00",
    "currency": "USD"
  },
  "confidence": 0.98,
  "fieldConfidence": {
    "Invoice Number": 0.99,
    "Amount Due": 0.98
  },
  "summary": "Vendor invoice from Acme Industrial Ltd totaling $1,350.00.",
  "createdAt": "2026-08-15T09:12:00.000Z"
}

Submit document for inference

POST /api/v1/infer

curl -X POST \
  -H "Authorization: Bearer dd_live_a8f92b7c4d1e0394..." \
  -F "file=@invoice.pdf" \
  https://app.dynamitedocs.com/api/v1/infer
{
  "id": "ext_99a8b7",
  "fileId": "file_77c2d1",
  "fileName": "invoice.pdf",
  "docType": "Invoice",
  "columns": ["Invoice #", "Vendor", "Due Date", "Total"],
  "rows": [
    ["INV-4410", "Apex Logistics", "2026-08-30", "$3,420.00"]
  ],
  "metadata": {
    "vendor": "Apex Logistics",
    "total": "3420.00",
    "currency": "USD"
  },
  "confidence": 0.97,
  "createdAt": "2026-08-22T14:10:00.000Z"
}

List extractions with pagination

GET /api/v1/extractions?page=1&limit=20&sort=created_at&order=desc

curl -X GET \
  -H "X-API-Key: dd_live_a8f92b7c4d1e0394..." \
  "https://app.dynamitedocs.com/api/v1/extractions?page=1&limit=20&doc_type=Invoice"
{
  "data": [
    {
      "id": "ext_01j5x9b3",
      "fileId": "file_8829f0",
      "fileName": "vendor_invoice_1092.pdf",
      "docType": "Invoice",
      "columns": ["Invoice Number", "Vendor", "Amount Due"],
      "rows": [["INV-1092", "Acme Industrial Ltd", "$1,350.00"]],
      "confidence": 0.98,
      "createdAt": "2026-08-15T09:12:00.000Z"
    }
  ],
  "pagination": {
    "page": 1,
    "limit": 20,
    "total": 48,
    "totalPages": 3,
    "hasMore": true
  }
}

Check monthly API usage

GET /api/v1/usage

curl -X GET \
  -H "Authorization: Bearer dd_live_a8f92b7c4d1e0394..." \
  https://app.dynamitedocs.com/api/v1/usage
{
  "usage": {
    "month": "2026-08",
    "calls": 42,
    "limit": 5000,
    "remaining": 4958,
    "unlimited": false
  }
}

What to validate before relying on it

Check the response status and usage headers on every call; a 401 means the key is invalid, while a 429 can mean the per-minute limit or monthly API allowance was reached.

The API is the difference between a tool and plumbing

An extraction that can only be exported by clicking a button is a report; an extraction that can be fetched by a token-authenticated GET is an integration. The API exists to move document data into the systems that already run the business, procurement and accounting, without a human re-keying anything. For a finance team ingesting vendor invoice rows into their accounting stack, or a support team pulling structured data from inbound emails into a ticket, the API is the whole integration.

Token design is where the security boundary matters. A key can list and export extractions, check its usage, and submit a document for inference. It cannot delete files, change settings, or touch billing. The server stores only the SHA-256 hash, so a database dump does not yield a usable credential. The full token is displayed exactly once at creation and can be revoked independently.

Callers can send a token in either the standard Authorization header, Authorization: Bearer dd_live_<key>, or the X-API-Key header. Both options work across all /api/v1 endpoints in curl, Postman, backend services and automation tools.

The export payload contains columns, rows, metadata and docType. CSV includes a UTF-8 BOM so Excel opens accented characters correctly. JSON carries the structure used by pipelines, while XLSX creates a workbook for spreadsheet work. The API is versioned under /api/v1, and the list endpoint paginates and filters in the database instead of returning the entire account history at once.

Every authenticated /api/v1 request is metered against the account’s monthly API-call allowance, and the gatekeeper rate-limits per key (default 60/min). The counter lives in the same usage ledger the app writes to, so the number a developer sees in GET /api/v1/usage is the number the server enforces. Revoking a key removes its hash and the very next request is rejected, which is why pasting a leaked key into a public repo is a one-click fix, not a fire drill.

Standalone inference fills the last gap. POST /api/v1/infer accepts a multipart document, returns the structured table, persists the extraction for the key owner, and triggers configured webhooks. Hosted runs use the same page caps and PE budget as the interactive app. BYOK on Free and Starter shares that PE budget; BYOK on Pro and Scale is unlimited. Every request still counts against the plan’s API-call allowance.

What the API does and does not

Does

  • Provides key-authenticated access to extractions and export formats via versioned /api/v1 endpoints with Bearer or X-API-Key auth.
  • Stores keys as SHA-256 hashes only and shows the raw secret exactly once at creation.
  • Rate-limits per key and meters every call against the monthly allowance, with a usage endpoint.
  • Enforces hosted page caps and the monthly PE budget on standalone inference. BYOK on Free and Starter is metered; BYOK on Pro and Scale is unlimited.

Does not

  • Does not expose file, settings, account, or billing mutations. Those actions require your session.
  • Does not allow API traffic to bypass rate limits, the monthly call allowance, or the processing budget.
  • Does not store your key in plaintext anywhere on the server.

Why the API belongs in the baseline

A token-authenticated export turns every extraction into a stable data source for the systems you already run, no human in the loop, no session cookie in a script.

Read-only, hashed tokens are safe to embed in pipelines and vendor tools, which is the only kind of credential you should be committing to config.

Standalone inference gives you programmatic document reads with the same caps and budget as the app, so automation cannot quietly game the pricing model.

Questions, answered

How do I create an API key?

From the workspace header under Webhooks & API, or directly via /app?modal=webhooks. The key is displayed once at creation and stored as a SHA-256 hash server-side.

How do I send my API key in requests?

Send your key in the Authorization header with Bearer scheme (Authorization: Bearer dd_live_<key>) or in the X-API-Key header (X-API-Key: dd_live_<key>). Both standard Bearer and custom header authentication are supported across all /api/v1 endpoints.

What can a key do?

Keys can list and export extractions, check usage, and submit documents through /api/v1/infer. File, settings, account, and billing changes stay behind your session.

Can I revoke a token?

Yes. Tokens can be revoked individually from Webhooks & API (/app?modal=webhooks), and revocation is immediate.

Does the API respect my plan limits?

Yes. Inference is rate-limited per user/IP. Successful hosted pages and Free or Starter BYOK pages draw from the monthly PE allowance. Pro or Scale BYOK pages cost 0 PE.

Related integrations

Document workflows it feeds

Related integration articles

Open the integrations workspace, no payment details.

Loading Dynamite Docs… This page is taking longer than expected. Reload page.