How to Use Your Own API Key for AI Document Extraction: Complete BYOK Guide
Dynamite Docs, 2026-08-12
The BYOK breakthrough: why pay 20x software markups on OCR?
Legacy Intelligent Document Processing software charges between $0.05 and $0.50 per page. Most vendors take raw cloud vision models, add a 1,000% token markup, and lock operations teams into rigid coordinate-based templates that break whenever a supplier changes their invoice header.
Modern multimodal large language models change that dynamic. State-of-the-art vision models read unstructured PDF invoices, crumpled receipts, and dense bank statements directly into structured JSON for fractions of a cent per page. Running a multi-page invoice through Gemini 2.0 Flash costs roughly $0.0002 in model tokens.
Bring Your Own Key, or BYOK, separates the document workspace from the AI inference engine. Dynamite Docs handles file intake, OCR text alignment, schema inference, human reconciliation, and spreadsheet exports. Your connected provider account executes the model inference and bills your card directly at wholesale cost.
On Dynamite Docs Pro ($79/month) and Ultra ($199/month) plans, BYOK extractions cost exactly 0 Processing Entitlements (PE). You can process thousands of document pages without software metering, paying only your raw model token costs directly to your selected provider. Free and Hobby plans can also connect provider keys within their monthly PE allowance.
- Wholesale token pricing: Pay the model provider raw API rates instead of inflated per-page software fees
- Zero template maintenance: Multimodal vision models infer line items and totals without coordinate boxes
- Unlimited BYOK on Pro and Ultra: Dynamite Docs charges 0 PE for own-key processing on paid tiers
- Review-first workspace: Inspect extracted fields side-by-side with original document pages before export to Excel or Google Sheets
Security architecture: how provider credentials and documents stay protected
Connecting an external API key requires strict credential hygiene. In Dynamite Docs, provider API keys never touch client-side browser storage after initial input. The server encrypts every key at rest using AES-256-GCM with server-side encryption keys managed via BYOK_ENCRYPTION_KEY.
The server decrypts your token in memory only when authenticating an HTTPS extraction request to the official provider endpoint. Keys are never returned in browser network payloads, exported in audit logs, or accessible to other workspace users.
A common concern in finance and legal operations is whether AI providers train on uploaded company ledgers. Standard commercial developer API terms from Google Cloud, Groq, Mistral, OpenAI, and Anthropic state that inputs and outputs sent through paid API endpoints are not used to train foundation models. This is fundamentally different from free consumer chat interfaces, which retain prompt history by default.
For organizations bound by strict regulatory constraints, Dynamite Docs allows policy routing. You can restrict sensitive supplier classes to European-hosted endpoints, enforce zero-retention policies, or route files exclusively to an on-premise Ollama instance.
- AES-256-GCM encryption: Provider keys are encrypted at rest with server-side keys and never sent back to browser clients
- Commercial API zero-training: Paid developer endpoints do not use document text or images to train commercial models
- Dedicated credential scoping: Generate separate keys with read/write model permissions and revoke them instantly if needed
- Ledger auditability: Track every job run, model version, token consumption, and human review sign-off
Google AI Studio: get a free Gemini API key for high-volume extraction
Google AI Studio provides developer access to the Gemini model family. For document processing, gemini-2.0-flash is one of the fastest and most cost-efficient multimodal models available. It parses complex invoice tables, handwritten dates, and tax breakdowns in under two seconds per page.
For multi-page contracts, annual audit reports, and dense loan binders, gemini-1.5-pro offers a 2,000,000 token context window. You can pass a 150-page statement batch in a single request without chunking or losing row relationships across page breaks.
Google AI Studio offers a free tier that allows up to 15 requests per minute (RPM) and 1,500 requests per day (RPD) for gemini-2.0-flash. For higher throughput, pay-as-you-go pricing starts at $0.10 per million input tokens, which amounts to roughly $0.00015 to $0.0003 per document page.
Follow these steps to generate your Google AI Studio API key:
- Step 1: Open the Google AI Studio API Key Console in your browser
- Step 2: Sign in with your Google account or Google Workspace administrator profile
- Step 3: Click the blue Create API key button in the top left corner
- Step 4: Select Create API key in new project (or link to an existing Google Cloud project with billing enabled)
- Step 5: Copy the generated string starting with
AIzaSy...and store it in your password manager - Step 6: In Dynamite Docs, open Settings > AI & Processing, select Google Gemini, paste your key, and click Verify & Save
Groq Cloud: sub-second vision extraction with Llama 3.2
Groq executes open-weights models on custom Language Processing Unit (LPU) silicon. Its architecture delivers processing speeds between 300 and 500 tokens per second, making it the fastest option for real-time receipt capture and high-velocity invoice intake.
When configured with Meta Llama 3.2 Vision models (llama-3.2-11b-vision-preview or llama-3.2-90b-vision-preview), Groq processes single-page invoices and expense slips in roughly 600 to 900 milliseconds. This speed allows operations teams to verify line items immediately upon upload rather than waiting in an asynchronous queue.
Groq provides a free developer tier with rate limits up to 30 requests per minute, followed by on-demand pricing that bills per million tokens with zero monthly platform minimums.
Follow these steps to generate your Groq API key:
- Step 1: Navigate to the Groq Cloud Console
- Step 2: Create a free account or log in with GitHub, Google, or email
- Step 3: Click API Keys in the left navigation sidebar
- Step 4: Click the Create API Key button and enter a label such as
dynamitedocs-vision - Step 5: Copy the secret key starting with
gsk_.... Groq displays this secret key only once upon creation - Step 6: In Dynamite Docs, open Settings > AI & Processing, select Groq, paste your key, and select
llama-3.2-11b-vision-previewas your active model
Mistral AI: European data sovereignty and complex multilingual tables
Mistral AI is based in France and operates dedicated European inference infrastructure. For organizations operating under European Union data protection regulations, Mistral provides direct GDPR compliance, European data residency, and clear subprocessor transparency.
Mistral developed the Pixtral model family (pixtral-12b and pixtral-large), which uses a dedicated vision encoder designed for spatial document reasoning. Pixtral handles multilingual invoices with mixed French, German, Spanish, and English labels without hallucinating column alignments or misinterpreting European comma decimal notations.
Mistral also offers specialized document OCR endpoints for reading native PDFs and scanned paperwork. Developer accounts receive pay-as-you-go pricing without long-term contract lock-ins.
Follow these steps to generate your Mistral API key:
- Step 1: Open the Mistral La Plateforme Console
- Step 2: Sign up for a developer account or sign in to your existing workspace
- Step 3: Select API Keys from the Workspace menu in the left navigation panel
- Step 4: Click Create new key, assign workspace permissions, and name the credential
dynamitedocs-pixtral - Step 5: Copy the secret token and store it securely in your password manager
- Step 6: In Dynamite Docs, open Settings > AI & Processing, select Mistral, paste your key, and test connectivity against
pixtral-12b
OpenAI and Anthropic: precision reasoning for dense audits and legal forms
When extracting data from dense 60-page credit agreements, government tax schedules, or detailed audit workpapers, reasoning accuracy takes precedence over raw inference speed. A single misread decimal point on a loan covenant or depreciation schedule creates hours of manual reconciliation.
OpenAI and Anthropic represent the industry benchmark for complex visual reasoning. OpenAI gpt-4o-mini offers an economical option for routine receipts at roughly $0.001 per page, while full gpt-4o handles degraded historical scans and multi-column tables. Anthropic claude-3-5-sonnet excels at interpreting footnotes, asterisks, handwritten margin notes, and crossed-out lines.
Both providers support strict JSON schema outputs, ensuring that dates, currencies, and line item arrays adhere exactly to your target spreadsheet columns.
Generate your credentials using these official console workflows:
- OpenAI setup: Go to the OpenAI Platform API Keys dashboard, click Create new secret key, select model permissions, and copy the
sk-proj-...token - Anthropic setup: Go to the Anthropic Console Settings, click Create Key, name it
dynamitedocs-claude, and copy thesk-ant-...key - Economical option: Use
gpt-4o-minifor high-volume standard invoices to keep token costs under $0.001 per document - Complex audit option: Select
claude-3-5-sonnetfor multi-tier balance sheets, complex debt schedules, and dense legal agreements
Local Ollama companion: 100% on-device extraction with zero cloud transmission
Certain files cannot be sent to third-party cloud infrastructure due to client non-disclosure agreements, HIPAA protected health information, attorney-client privilege, or internal security controls. For these scenarios, Dynamite Docs includes an integrated local companion that routes inference to an Ollama instance running directly on your computer.
Document images and extracted fields are processed over a local loopback connection on your workstation (http://127.0.0.1:11434 and companion port 8756). No document bytes or text tokens travel over the external internet. When your computer is offline or disconnected from Wi-Fi, extraction continues uninterrupted.
Local processing requires adequate hardware. Modern Apple Silicon Macs (M1/M2/M3/M4 with 16GB or more unified memory) or Windows/Linux PCs with an NVIDIA RTX GPU can run llama3.2-vision:11b or minicpm-v at practical extraction speeds of two to four seconds per page.
Follow these steps to configure private on-device extraction:
- Step 1: Download and install the Ollama runtime from Ollama Official Download
- Step 2: Open your terminal and pull a multimodal vision model: run
ollama run llama3.2-visionin your command line - Step 3: For workstations with limited GPU memory, pull the compact model: run
ollama run minicpm-v - Step 4: Download the Dynamite Docs Windows companion or verify Ollama is serving on port 11434
- Step 5: In Dynamite Docs, navigate to Settings > AI & Processing, click the Local Ollama tab, and confirm the connection status indicator is green
How to connect and verify your API key inside Dynamite Docs in 60 seconds
Connecting an external API key in Dynamite Docs takes less than one minute. Once connected, your credentials persist securely and apply to all manual uploads, folder imports, and batch extraction jobs.
Start by opening Dynamite Docs Settings using the gear icon in the workspace navigation. Click the AI & Processing tab. Choose your provider from the dropdown list, paste your secret key into the credential field, and click Verify & Save Key.
Dynamite Docs sends an encrypted validation ping directly to the provider endpoint. This test verifies that your token is active, checks your account permissions, and discovers the vision models available to your tier. Once verified, the provider card displays a green connection badge with your active model selection.
Always test your setup with a representative sample file before launching a large batch. Upload an invoice with multiple line items, check that the columns extract cleanly into the table grid, and verify that confidence scores highlight any low-contrast values.
- Open settings: Navigate to Settings > AI & Processing in your Dynamite Docs workspace
- Select provider: Choose from Google Gemini, Groq, Mistral, OpenAI, Anthropic, OpenRouter, or Custom Endpoint
- Verify connection: Click Verify & Save Key to perform a live permission and quota health check
- Configure fallbacks: Set a secondary model route to handle temporary rate limits or provider downtime automatically
Troubleshooting rate limits, token limits, and multi-page batch runs
When processing large document batches with your own API keys, you may encounter provider-specific rate limits or token constraints. Understanding these constraints prevents interrupted batch jobs and keeps extraction workflows running smoothly.
HTTP 429 errors indicate that your job has exceeded the provider requests-per-minute (RPM) or tokens-per-minute (TPM) quota. Provider free tiers typically restrict throughput to 15 or 30 RPM. Dynamite Docs includes built-in exponential backoff retry mechanisms to handle brief rate spikes. For production batches exceeding 200 pages daily, upgrading to a pay-as-you-go card on Google AI Studio or Groq raises limits to hundreds of requests per minute with minimal cost.
Document image resolution directly affects vision token usage. Multimodal models divide document pages into 512x512 pixel tiles. A standard 300 DPI invoice scan uses approximately 1,000 to 1,600 vision tokens. Avoid uploading uncompressed 1200 DPI scanner TIFF files, as they consume excessive tokens without improving character recognition.
When switching between different models like Gemini, Claude, and Llama, Dynamite Docs schema locking ensures that column names remain identical. The application standardizes field outputs so that your downstream Excel, CSV, or Google Sheets formulas never break.
- HTTP 429 management: Enable automatic retry backoff and size batch queues to fit provider rate limits
- Resolution sizing: Target 200 to 300 DPI for scans to balance character clarity with token efficiency
- Token budgeting: Estimate roughly 1,200 tokens per invoice page when calculating wholesale API costs
- Schema locking: Maintain consistent table column headers across different AI model versions and providers
Frequently asked questions about BYOK document extraction
Common questions from finance, operations, and IT teams implementing BYOK document extraction workflows.
- Are AI provider API keys free? Several providers offer free tiers. Google AI Studio provides free access to Gemini 2.0 Flash up to 15 RPM, and Groq offers free developer tiers for Llama 3.2 Vision. For production workloads, pay-as-you-go rates cost fractions of a cent per page ($0.00015 to $0.002), far lower than traditional OCR subscriptions.
- Does Dynamite Docs charge Processing Entitlements (PE) when I use my own key? On the Pro ($79/month) and Ultra ($199/month) plans, BYOK document extraction costs exactly 0 PE. You can process unlimited pages with your own key without platform metering. On Free and Hobby plans, BYOK extractions draw from your shared monthly PE allowance.
- Will the AI provider train their models on my invoices, receipts, or contracts? When using paid commercial API keys from Google Cloud, OpenAI, Anthropic, Mistral, and Groq, data sent through API endpoints is not used to train foundation models under their standard commercial terms of service. For complete air-gapped security, use the local Ollama companion with our local document AI guide.
- Do I need to write code or prompts to use my API key in Dynamite Docs? No code or prompt engineering is required. Dynamite Docs handles document parsing, prompt construction, coordinate extraction, JSON validation, and export formatting. You simply paste your API key once in settings and use the visual workspace.
- Can BYOK vision models extract handwritten receipts and blurry smartphone photos? Yes. Modern multimodal models like Gemini 2.0 Flash, Claude 3.5 Sonnet, and GPT-4o excel at reading handwriting, angled mobile phone photos, and degraded paper receipts that cause traditional OCR engines to fail.
- What happens when a provider encounters downtime or rate limits? Dynamite Docs allows you to configure primary and secondary provider routes. If your primary provider returns an error or rate limit, the system can automatically retry or route extraction through your approved fallback provider.
Start with a pilot test on your most difficult document layout
Bring Your Own Key architecture provides total control over model selection, wholesale pricing, and data residency. It eliminates the 20x markups charged by legacy document extraction vendors while unlocking the power of modern multimodal vision models.
Before rolling out batch extraction across your organization, run a small pilot test. Generate a free key from Google AI Studio or Groq Console, connect it in Dynamite Docs Settings, and process five of your most difficult vendor invoices. Compare the extracted line items against the original PDF and verify the results before expanding to full production volume.
Sources checked for this note
- Google AI Studio API Key Management
- Groq Cloud Developer Console
- Mistral AI La Plateforme Console
- OpenAI Developer Platform API Keys
- Anthropic Developer Console Settings
- Ollama Official Runtime Download
- OpenRouter API Key Management
- NIST SP 800-38D: Recommendation for Block Cipher Modes of Operation (Galois/Counter Mode)
Related workflow: Accounting document automation and data extraction.
Keep reading
Try it yourself. Upload a PDF, scan, or image and let Dynamite Docs infer the schema.