All India workflows
GST invoice OCR — read scanned and photographed invoices
A GST invoice that reaches you as a scan, a fax, or a phone photo has no text layer. You cannot copy the GSTIN, search the invoice number, or select the HSN/SAC code, because the document is a picture of a page, and someone has to read it and retype it. That is the OCR problem Dynamite Docs solves: scanned, faxed, and photographed GST invoices are read by a vision model into the same confidence-scored rows as a clean digital PDF, with GSTIN, HSN/SAC, taxable value, and the CGST/SGST/IGST split in place. Clean digital PDFs skip the vision step entirely and parse deterministically for free.
Try this workflow free or view pricing.
Paper and photos are the slow part of GST bookkeeping
- Scans and faxes carry no text layer, so every GSTIN, invoice number, and amount on them has to be read by a person and keyed in by hand, with no copy-and-paste to fall back on.
- Phone photos arrive at bad angles with glare, shadows, and clipped edges; crumpled paper and staples bend the lines a template-based tool depends on.
- The fields that matter most, the 15-character GSTIN, the HSN/SAC code, and the CGST/SGST/IGST split, are exactly the ones most likely to be faded, smudged, or cut off at the photo edge.
- Plain OCR gives you a wall of text. The bookkeeping task is a table of invoice number, vendor, item lines, taxable value, and taxes, and that table is what actually needs to come out.
How GST invoice OCR works
01 — Upload the scans and photos
Drop in the PDFs, JPGs, and WhatsApp-forwarded photos as they arrive. There is no prep, no deskewing step, and no template to line the document up against.
02 — Choose the vision model
Scanned and photographed invoices are read by the vision model you pick. Bring your own key from Gemini, Mistral, OpenAI, Anthropic, Groq, Cloudflare, or OpenRouter to control quality and per-page cost, or keep the default hosted model.
03 — Review by confidence
Each field carries a confidence score, so a smudged GSTIN or a glare-covered amount is flagged instead of silently guessed. Corrections you make become patterns the next batch reuses.
04 — Export clean rows
The structured result is the same as any other extraction: GSTIN, invoice number and date, HSN/SAC, taxable value, CGST, SGST, IGST, and totals, exportable to Excel, CSV, or JSON.
Document types this workflow handles
- Scanned GST invoice PDFs — no text layer, read by the vision model
- Faxed and photocopied invoices — faded and low-contrast originals
- WhatsApp and email photos — angles, glare, and shadows handled
- Thermal-paper purchase bills — low-contrast prints that fade
What gets extracted from a scanned GST invoice
The same field set as a digital PDF, each value confidence-scored so the weak reads stand out:
- GSTIN (supplier) — 07AAECS4458H1ZF
- Invoice number — GT-8821
- Invoice date — 2026-08-05
- Vendor / customer — Vardhan Packaging
- Place of supply — Delhi (07)
- HSN / SAC code — 4819
- Item & quantity — Corrugated boxes × 800
- Taxable value — ₹36,000.00
- CGST (rate · amount) — 9% · ₹3,240.00
- Invoice total (incl. tax) — ₹42,480.00
A realistic example
A batch of photographed and scanned invoices extracts to the same rows as digital PDFs, flagged by confidence where the scan is weak:
| Invoice # | Date | Vendor | GSTIN | HSN/SAC | Item | Taxable value | CGST | SGST | IGST | Total |
| GT-8821 | 2026-08-05 | Vardhan Packaging | 07AAECS4458H1ZF | 4819 | Corrugated boxes | ₹36,000 | ₹3,240 | ₹3,240 | — | ₹42,480 |
| KT-5512 | 2026-08-07 | Ganesh Industries | 29AABCC3456C1ZD | 8467 | Rotary hammer | ₹1,86,000 | — | — | ₹33,480 | ₹2,19,480 |
| TH-0916 | 2026-08-08 | Prabhat Hardware | 27AAEHP5621L1ZV | 7323 | Steel pipes | ₹28,000 | ₹2,520 | ₹2,520 | — | ₹33,040 |
What OCR has to survive to be useful
The difference between OCR that works and OCR that creates more work is whether it hands you a table. A raw text read of a photographed invoice gives you lines of text in reading order, which a bookkeeper still has to interpret. Dynamite Docs returns the structure: GSTIN, invoice number, date, vendor, HSN/SAC, item lines, taxable value, the CGST/SGST/IGST split, and the total as separate fields, each with a confidence score. You review the flagged values instead of reading the whole document again.
The real-world test is the thermal receipt and the phone photo. Thermal paper fades within months, and photographed invoices come in at odd angles with glare across the tax block. The vision model is chosen for exactly this, and because you control the model on BYOK you can trade cost against quality per batch rather than accepting a fixed pipeline. Deterministic parsing of clean digital PDFs stays the free path, so the tool does not bill vision for documents that do not need it.
The structure is what makes OCR useful for GST work. A GSTIN either matches the 15-character pattern or it gets flagged; CGST and SGST pair up on intra-state invoices and IGST on inter-state ones, so a split that contradicts the place of supply stands out. That review step stays with your team, and when you prepare GSTR-1 you work from the rows in your own software. Consult the GST portal and your tax advisor for current rules on anything the extraction surfaces.
Why the vision model is your choice, not a fixed pipeline
OCR quality is a model property, not a setting. Thermal prints, low-contrast photocopies, and bad-angle photos read differently on different vision models, which is why the reader is your choice rather than a fixed pipeline: bring your own key, pay your provider’s rate with zero markup, and on Pro and Ultra process without the OCR ever counting against your monthly allowance.
The batch never needs one path. The clean digital PDFs that make up most of a month’s invoices parse deterministically for free, while only the true scans and photos spend vision credits on the model you pick. Cost follows the document, not the folder.
The flagged low-confidence reads, a smudged GSTIN, a glare-covered amount, are the ones worth a glance, and each correction becomes a pattern the next batch from the same sender reuses. The structured rows, review flags, and corrections export cleanly to Excel, CSV, or JSON.
Questions, answered
Can it read a faded or crumpled scan?
Scans and photos are read by the vision model you choose, which is what handles faded thermal prints, low-contrast photocopies, and crumpled paper. Fields that come out with low confidence are flagged for your review rather than passed through silently.
What happens with a clean digital PDF?
It is parsed deterministically from its text layer, no AI call, identical results every time, at no cost. Vision is only used where the document actually needs it.
Which vision models can I use?
Bring your own key from Gemini, Mistral, OpenAI, Anthropic, Groq, Cloudflare, or OpenRouter, or run locally with Ollama. BYOK pays your provider’s rate with no markup, and on Pro and Ultra it is unlimited and never counts against your monthly allowance.
Does OCR handle WhatsApp photos at bad angles?
Yes. Bad angles, glare, and shadows are exactly what the vision model is for. The output is still the standard field set, GSTIN, invoice number, HSN/SAC, taxable value, and the tax split, each confidence-scored and reviewable.
Related document workflows
Open app, no card needed.