Image OCR & Text Extraction Tool | Export Tables and Forms to CSV/TSV

English OCR processed on your devicePrinted text, receipts, forms, and tablesProcess multiple images in sequenceNo sign-up or installation

Choose image to read

Adjust image processing, text direction, and text layout to suit your use case.

Choose image

Drag image hereJPEG, PNG, WebP, BMP, up to 20 images.
Prioritize the device’s rear camera to photograph the document on the spot.
No image selected yet.

Text recognition settings

Adjust rotation, tilt, and cropping
Margin to crop (percentage from each edge of the image)
Clean up recognized text

Image and OCR Text Result

Select an image to preview the corrected state.

0 / 0
Image will appear hereThe more horizontal the text and the fewer shadows and margins, the easier it is to recognize.
Waiting0%

Images and recognition results are not sent to a server. The OCR engine and dictionaries are also loaded from WiseChecker.

What this image OCR tool can do

Extract printed English text and numbers from scans, phone photos, receipts, forms, labels, and tables. Review the result, copy it, or save TXT, CSV, and TSV files.

Images and recognition results remain in your browser. The OCR engine and English dictionary are served from WiseChecker rather than an external OCR service.

English printed-text OCR

Recognize paragraphs, labels, dates, amounts, and identifiers.

Document image cleanup

Adjust grayscale, contrast, thresholding, inversion, and crop.

Review low-confidence text

Jump to words that need comparison with the original image.

Export tables

Estimate columns and save CSV or TSV for further cleanup.

Steps to transcribe text from an image

  1. Select image typeChoose the closest use case: documents, invoices/receipts, tables/forms, or vertical text.
  2. Load imageOn PC, you can use image selection, drag-and-drop, and pasting. On compatible smartphones and tablets, you can launch the device camera from “Capture Document with Camera.”
  3. Review corrected imageAdjust orientation, tilt, margins, and text darkness to make the text easy to read.
  4. Run OCRProcesses a single image or all loaded images sequentially. The first run requires time to load the OCR engine and dictionaries.
  5. Cross-check verification candidates with the original imageFocus on checking words with confidence below 70, amounts, dates, part numbers, and proper nouns.
  6. Save in a format that fits your useSave text as TXT and tables as CSV or TSV. You can also edit on screen and then copy.

Settings for reading invoices, receipts, and forms

Selecting “Invoices and Receipts” applies binarization that reduces uneven brightness and paper shadows, reading the content as a continuous region. If dates, slip numbers, registration numbers, amounts, tax rates, and product names are widely separated, switching “Text Layout” to “Scattered Text and Items” may help.

In “Tables and Forms,” spaces between words are preserved, and column breaks are estimated from the horizontal positions of recognized words. The table tab displays tab-separated values, ready to paste into Excel. For CSV saving, fields containing commas, such as product names or amounts, are enclosed in quotes.

Image typeChoose use case firstReview settingsCharacters to check specifically
InvoiceInvoices and receiptsShadow correction, distant text and fieldsInvoice number, registration number, amount, tax rate
ReceiptInvoices and receiptsCrop margins, correct tiltDate, quantity, unit price, total, store name
Application forms and checklistsTables & formsKeep blanks, convert to black and whiteCheckboxes, entered values, units, margin notes
Notices and minutesDocuments and noticesAuto paragraph detection, gray correctionProper nouns, dates, URLs, email addresses

How to take photos to improve OCR accuracy

OCR accuracy depends on character size, focus, skew, shadows, background, and compression. When using a smartphone, hold it parallel to the document and position it so all four corners are within the frame. If overhead lighting reflects strongly off the paper, slightly change the angle of the device or light.

  • Capture at a resolution where character outlines are visible, and use the original image rather than one recompressed via social media or chat.
  • For photos where the edges of the paper darken, choose “Reduce shadows and convert to black and white.”
  • Remove non-document surfaces like desks and backgrounds by cropping margins.
  • Use the rotate left/right buttons for 90-degree rotations, and the tilt correction for slight angles.
  • Increases contrast for faint print. For blurred characters, binarization can be counterproductive, so also compare with grayscale correction.

Small images are automatically enlarged. Enlargement does not restore information absent from the original, but it standardizes the size so the OCR engine can analyze character outlines more easily. If characters are lost due to low resolution, blur, or overexposure, retaking the photo is the most reliable option.

Difference between gray correction, binarization, and shadow correction

Image correctionSuitable imagesOperationIf it doesn’t fit
Grayscale and contrast correctionScans, printed documents, faint textConverts color to brightness and expands dark and light areas.Strong background shadows may be recognized as text.
Reduce shadows and convert to black and whiteSmartphone photos, receiptsSeparates text into black and white by comparing with surrounding brightness.Shading, photos, and thin rules may disappear.
BinarizeForms with clear contrastDetermines threshold values from the entire image and separates black and white.Uneven lighting may cause some characters to be missing.
No correctionAlready clear scanOnly rotate, crop, and enlarge as needed.Colored backgrounds and faint print are harder to recognize.

The preview shows the actual image passed to OCR. If characters appear clipped after correction, do not proceed; switch to a different correction. Image adjustments are retained per selected image.

Choose the recognition language and layout

This version uses the English training dictionary. Choose automatic layout for most documents, single block for a continuous paragraph, single line for labels, or sparse text for scattered fields.

Mixed fonts, handwriting, decorative lettering, and very small text reduce accuracy. Crop to one reading area and compare important values with the original.

How tables and forms are converted to Excel-ready CSV/TSV.

Using the left position, width, height, and line number of each word returned by the OCR engine, it estimates column boundaries where large gaps occur within a line. It does not restore cells from ruled lines themselves. Therefore, empty cells, merged cells, multi-line cells, and checkboxes may not match the original table structure.

CSV is saved as a file, while TSV is suited for pasting into Excel or spreadsheets. If columns are misaligned, you can edit the tabs in the table tab before saving. When importing amounts or dates into a spreadsheet, verify that thousands separators, minus signs, decimal points, and era names match the original image.

Order to verify after form OCR

  1. Whether row and column alignment is preserved
  2. Check that 0 and O, 1 and I/l, 5 and S, and 8 and B are not swapped.
  3. Check that commas, decimal points, hyphens, parentheses, and currency symbols remain.
  4. Whether the total matches the sum of line items
  5. Check that blank fields and checkboxes are not mistakenly recognized as text.

How to read confidence and candidates

Average confidence is a reference value (0–100) averaging how confident the OCR engine is in each word. This tool lists words with confidence below 70 as “confirmation candidates.” If confirmation candidate display is enabled, you can also see their positions highlighted with red boxes on the image.

Even with high confidence, words can be misrecognized as plausible alternatives in context. Especially for amounts, account numbers, phone numbers, addresses, dates, invoice numbers, registration numbers, product model numbers, and personal or company names, do not rely solely on confidence—always verify against the original. Conversely, decorative characters or proper nouns may be correctly recognized but have low confidence.

When transcribing multiple images at once

You can select up to 20 images, 20 MB each, and 150 MB total. “Transcribe All” processes images one by one in the on-screen order. OCR uses your device’s CPU and memory, so it runs sequentially rather than in parallel to reduce load.

Results are retained per image. Selecting a different image in the left list switches to that image’s correction state, transcription result, table format, and confirmation candidates. TXT saving combines all images into one file with filenames as headings; CSV/TSV saving includes page breaks.

Many high-resolution images increase browser memory usage. If performance is slow, split into batches of about five images and close other heavy tabs. Pressing the stop button terminates the current OCR process and keeps completed results.

Clean the OCR result without losing structure

Trim line edges, collapse repeated spaces, remove blank lines, or join wrapped lines only after checking the original layout. Forms and tables often use spacing to show columns, so keep an untouched copy before aggressive cleanup.

OCR runs locally in this browser

The selected image is decoded, corrected, and recognized on your device. WiseChecker does not upload the image or recognized text to an OCR processing server.

Tesseract.js, its WebAssembly core, and the English training data are served from WiseChecker. Closing or reloading the page clears the current working list.

What OCR cannot read accurately

This tool targets printed text. It does not accurately restore handwritten characters, extremely stylized fonts, watermarks blending into the background, low-resolution surveillance camera images, or characters lost to blur or mosaics. It does not use generative AI to guess and fill in information unreadable from the image.

TargetCommon issuesAction
Handwritten textCharacter shapes may not match the training data, causing major misrecognition.Do not rely on the results; have a person verify the original and enter data.
Complex tablesMerged cells, blank fields, ruled lines, and multi-level headers cannot be restored.Recognize in ranges and arrange the table on the Excel side.
Seals and stampsCharacters overlapping the body text may be missing.Use the original without a seal imprint or another copy.
Angled photoDistorts into a trapezoid, making rows and character widths uneven.Retake the photo with the document parallel to the camera.
PDF within imagePDF files cannot be selected in this input field.Load the required pages as PNG or JPEG.

Frequently asked questions about image OCR and form transcription

Are images and recognition results sent externally?

Nothing is transmitted. Image processing, OCR, and result generation occur in your browser, and the engine and dictionaries are loaded from within WiseChecker.

How many images can I transcribe for free?

Up to 20 images at once, 20 MB each, and 150 MB total. There is no limit on the number of uses, but to avoid device load, split large batches.

What is the difference between “Photograph a document” and “Select an image”?

“Select Images” loads images saved on your device. “Capture Document with Camera” appears only on compatible smartphones and tablets and is for capturing documents with the rear camera. On PC, camera capture is not available and behaves like image selection, so the capture button is not shown.

Why does the first start take time?

This is to load the OCR engine and dictionaries for the selected language onto your device. In the same browser, they are cached, which may shorten subsequent load times.

Can you read handwritten application forms and surveys?

Designed for printed text, so handwriting accuracy cannot be guaranteed. Only field names are recognized; fill-in content must be verified by a human.

Can the table be converted to the same Excel layout as the original?

Columns are estimated from character spacing, but ruled lines, merged cells, blank cells, column widths, and formatting are not restored. Open the CSV/TSV in Excel, compare with the original, and adjust.

Can amounts and account numbers be used directly in business?

OCR results may contain misrecognition. For critical information such as amounts, account numbers, dates, and registration numbers, verify against the original regardless of confidence.

Why can’t I select HEIC or PDF?

Limited to JPEG, PNG, WebP, and BMP, which can be reliably converted to images across browsers. For iPhone photos, capture or export in JPEG-compatible format, or convert to an image before selecting.

What to do if the image becomes black or text disappears.

Revert correction to “Grayscale and Contrast Correction” or “No Correction,” and check that invert colors is not mistakenly enabled.

Can I stop the process midway?

You can press “Stop Processing” during processing. It terminates the current OCR engine and keeps only the results for completed images.

Related free tools