Image OCR

Extract structured rows from photos, screenshots, and raw image files

Turn a document photo or screenshot into structured data while keeping labels, columns and source regions connected.

An image has no text layer to rescue a weak result. Focus, resolution, light, perspective and compression determine whether the system can distinguish a decimal point from dust or keep a value under the right column.

For document work, image OCR should return fields and rows rather than only a transcript. Use the original image with the whole page visible, and keep source regions with the result so a reviewer can trace a doubtful value back to the photo.

Where image ocr fits

  • Phone photos of paper documents
  • Screenshots that contain forms or tables
  • PNG, JPG, JPEG, and WebP scans

How image ocr works

  1. Check image quality: Measure resolution, blur, glare, orientation, cropping, and whether small text is still legible.
  2. Correct geometry: Deskew the page and correct perspective so rows and labels line up.
  3. Read by region: Recognize text while preserving its page position, grouping, and likely field role.
  4. Return reviewable data: Provide fields or rows with the source image available for visual comparison.

What the structured result can contain

Field or structureExample
Text regionsParagraph, label, value, caption
Key-value fieldsOrder number: PO-7714
Table cellsRow 4, amount 128.50
Reading orderHeader before body and footnotes
CoordinatesSource region for review

A better source image means less correction work

Image preparation often improves the result more than retrying the same poor file with another model.

  • Use the original image and highest practical resolution.
  • Photograph the page straight on with even light and all corners visible.
  • Avoid fingers, shadows, folds, and highlights over important text.
  • Do not over-sharpen or apply filters that erase punctuation and thin strokes.
  • Split contact sheets or collages into one document image per file when possible.

Prepare image OCR inputs without erasing useful evidence

Use the original photo or scan whenever possible. Repeated compression can blur punctuation, thin type and grid lines even when the image still looks acceptable on a phone. Keep all page edges visible so perspective can be corrected and the model can tell where a form or table begins.

Improve capture before adding aggressive filters. Even light, a camera held parallel to the page and sharp focus usually provide better evidence than heavy contrast or sharpening. A filter that removes a faint decimal point may make the page look cleaner while making the extracted amount wrong.

Choose the output around the document. A screenshot of a table needs rows and columns. A photographed form needs labels and values. Retain dimensions and source coordinates when the review interface will draw highlights over the image, then route unreadable regions to a person.

OCR API response design

Declare supported MIME types and size limits, and reject empty or unsupported files before inference.

Preserve image dimensions and return normalized coordinates if another interface will draw highlights over the source.

Read the document extraction API guide.

Common uses

Mobile capture

Extract fields from a document photographed outside the office.

Screenshot recovery

Turn a table or form trapped in an image into editable data.

Mixed upload intake

Accept [image documents alongside PDFs](/ai-ocr-software-for-document-data/extract-scanned-pdf-data-using-ai) in one extraction workflow.

Free tools for this document

Image OCR questions

Which image format is best for OCR?

A clear PNG or high-quality JPEG usually works well. Resolution, focus, contrast, and compression matter more than the file extension.

Can OCR read text from a screenshot?

Yes, if the text is large enough and not heavily compressed. Screenshots often lose page context, so keep the whole form or table visible.

How can I improve OCR on a phone photo?

Use even lighting, hold the camera parallel to the page, include all four corners, tap to focus, and upload the original photo.

Related OCR guides

Related Workflows

Loading Dynamite Docs… This page is taking longer than expected. Reload page.