Local document AI

Run private document extraction models locally on your own machine

Extract messy PDF tables into clean Excel spreadsheets on your local GPU using Ollama, then review cell confidence in the workspace.

The Dynamite Docs local companion sends document bytes used for model inference directly from the browser to a service bound to loopback on your workstation. The companion prepares the request for Ollama and returns a validated extraction to the browser without calling an external model API.

Local LLM support is available on Starter, Pro and Scale. Model inference runs on your hardware, while authentication, token issuance, account state and workspace services remain hosted.

Set up local extraction. Local Ollama requires Starter, Pro or Scale, a running companion, and a compatible vision model.

What runs where

  • Browser workspace: Starts the extraction and displays the editable result.
  • Signed local companion: Accepts authenticated requests on loopback port 8756.
  • Ollama: Runs the selected local model through its OpenAI-compatible API.
  • Hosted account service: Issues short-lived tokens and keeps plan and workspace state.

Local models return the same review shape

The companion validates the model response before returning document type, schema, rows and confidence data to the workspace.

  • Document type: invoice, statement, form or report
  • Suggested schema: Named columns inferred from the file
  • Rows: Structured records for review
  • Metadata: Document-level context and fields
  • Confidence: Result and field-level review signals
  • Source content: Prepared document context for comparison

From file to reviewed result

  1. 1. Install Ollama and a model. Choose a model that supports vision and tabular parsing, such as qwen2.5-vl:7b, and fits available VRAM.
  2. 2. Install the Windows companion. Install the companion tray app, then leave it running while you use local models. The companion listens on loopback port 8756 and Ollama defaults to 11434. Download the Windows companion
  3. 3. Connect from AI Settings. The hosted service issues a short-lived signed token and discovers models available through Ollama.
  4. 4. Extract and review locally. The browser sends the file to the companion, which prepares pages, calls Ollama and returns a validated result.

Select local models in the same AI panel

The workspace lists the chosen local model alongside hosted and own-key options. Local-only routing can prevent a document from falling back to an external provider.

The actual Dynamite Docs library import menu, with upload and Google Drive options.
The actual AI processing dialog showing the searchable model list and provider filters.
A purchase order beside its extracted text in the Text Editor.
A purchase order beside its structured details in the Table Editor.
A purchase order beside its extracted line items in the Table Editor.
A purchase order beside the Table Editor Totals tab, showing tax and the final total.
Purchase order text beside extracted document fields and confidence indicators.
The actual Export to Google Drive dialog with format, scope, file name and folder options.
Dynamite Docs routing rules showing local-only and sensitive-document model policies
The real routing rules panel used to restrict model selection.

The request stays on a short local path

The companion only listens on loopback by default. Every route requires a valid token, and browser origins are restricted to the official app, configured origins and local development.

StepEndpointResult
1. TokenHosted /api/local/tokenShort-lived signed local token
2. ModelsCompanion /v1/modelsOllama models available on the machine
3. ExtractCompanion /v1/extractValidated schema, rows and confidence
4. ReviewBrowser workspaceEditable extraction and export choices

Local does not mean effortless or identical

Confidence

The workspace still shows confidence and [editable results](/docs/reviewing-extractions-in-data-studio). Review money, identifiers and low-quality scans as you would with any provider.

Accuracy

Local model quality varies by model, quantization, document type and hardware. A smaller model may be faster but miss dense tables or small text.

Failure state

Extraction stops when Ollama is not running, the model is missing, the signed token expires, the companion cannot be reached or the model returns an invalid result.

Product boundary

The companion has a 150 MB upload limit, a default 120-second timeout and batches scanned PDF images within the model image limit. Hardware still sets practical speed and page limits.

Local and BYOK solve different problems

Use local inference when model calls should run on the machine. Use BYOK when an external provider account should be selected and billed directly.

OptionModel runsPlan access
HostedDynamite Docs selected providerFree, Starter, Pro and Scale
Local OllamaUser computer through the companionStarter, Pro and Scale
BYOKExternal provider account chosen by userEvery plan

The processing boundary is specific

Document bytes used for local inference move from the browser to the loopback companion and then to Ollama on the same machine. The companion does not expose an unauthenticated public extraction endpoint.

Account, plan, workspace and token issuance still use the hosted Dynamite Docs service. Teams that need a fully self-hosted application should not treat the local companion as that product.

Questions about local document ai

Which plans include local Ollama?

Starter, Pro and Scale include the local companion. Free does not. This entitlement is separate from BYOK.

Does the local companion work offline?

Model inference runs locally, but the current product still uses the hosted account and workspace service for authentication, token issuance and account state.

Which Ollama model should I use?

Choose a model that supports document images (vision vs. text) and fits your hardware. Test it on representative files, then compare missing fields, table structure and correction time before using it for a batch.

Can I run document extraction locally with Ollama without exposing data to external APIs?

The model-inference request stays on the local loopback path between the browser, companion and Ollama. Account, token and workspace services remain hosted, so review the full storage and retention path before treating the workflow as self-hosted.

Choose the next document task

Related Workflows

Set up local extraction or compare processing, storage and integration limits.

Loading Dynamite Docs… This page is taking longer than expected. Reload page.