OCR PDF

Make scans searchable while keeping the original page appearance.

Choose PDF Files

Drop files here or use the button below.

Up to 100 MB total · Files stay on your device

Table of Content12 Sections

When Does a PDF Need OCR?

Add a searchable text layer to a scanned PDF with Tesseract running in your browser. Choose the language that matches the source; English is selected by default. Supported readers can then search and select recognized words while the scanned page keeps its original appearance.

By default, the operation keeps pages that already contain text. You can also recognize every selected page; this adds a new layer and may duplicate any existing text. OCR accuracy depends on focus, contrast, language and page alignment. Small print, handwriting, unusual typefaces and noisy scans need careful review rather than automatic trust.

OCR PDF Features

Searchable PDF or TXT

Download recognized wording as plain text or add a text layer over the retained scan appearance in a PDF.

Up to 20 Selected Pages

The download contains only selected pages. Choose up to 20 pages; leave Pages blank to include the whole document only when it has 20 pages or fewer.

Recognition Controls

Choose the source language, 200 or 300 DPI, a recognition rotation, and automatic, single-block or sparse-text page layout. The supported languages are listed in the FAQ below.

Keep Existing Text by Default

Pages already containing text can be retained. Force recognition only when you have considered possible duplicate text layers.

Why Add a Searchable Text Layer to a Scan?

A scan stores the page as a picture, so ordinary text search cannot find the words printed on it. OCR in the selected language can make a typed scan searchable and provide wording for copy-and-paste or later text extraction. It is especially useful for archives where finding a known name or phrase matters more than redesigning the page. The recognized layer can contain mistakes even when the scan looks clear. Test search and selection, then compare names, references and numbers with the page image before using extracted wording.

Who Uses OCR on PDF Scans?

Library and Records Staff

Make a short typed record searchable by title, name or reference number. Keep the original scan and test those search terms in the downloaded copy, including any easily confused characters.

Students with Scanned Handouts

Make typed course notes searchable before building study notes. Handwritten annotations may remain unreliable, so check quoted passages against the scan.

Routine Personnel Administrators

Extract wording from an approved scanned memo for an internal index. Compare dates, staff names and reference numbers carefully because a single recognized character can change an identifier.

How Do You Make a Scanned PDF Searchable?

  1. Choose an unencrypted scan and select no more than 20 pages. The PDF download contains only that selection. Start with a short sample if you have not used this type of scan before.
  2. Select the source language in Text language. Choose Searchable PDF or Plain text, set recognition resolution, and select a rotation or page layout if the scan needs it.
  3. Keep existing text by default, or deliberately choose Recognize every selected page. Start OCR and wait for recognition to finish.
  4. Download the result. Search a known phrase, copy a line and compare spelling and numbers with the original scan.

Choose the Source File

Choose a scan in a supported language and select up to 20 pages.

Select the screenshot to enlarge it.
OCR PDF, step 1: Choose a scan in a supported language and select up to 20 pages.

Review the Tool Settings

Choose the source language, OCR output, resolution, layout and recognition orientation.

Select the screenshot to enlarge it.
OCR PDF, step 2: Choose the source language, OCR output, resolution, layout and recognition orientation.

Download and Check the Result

Download and test recognized text against the original scan.

Select the screenshot to enlarge it.
OCR PDF, step 3: Download and test recognized text against the original scan.

Screenshots show this toolkit using harmless sample files. Your file name, page count and settings may differ.

An Image-Only Checklist Made Searchable

The source page is an image with no extractable text. OCR adds a text layer while keeping the page’s visible appearance. Both images therefore look the same; the downloadable result also contains the recognized words.

Before: Image-Only Page

Text extraction from the source returned zero characters.

Actual image-only checklist before OCR with no selectable text layer.

After: Same Page, Searchable Text

The result contains all six sample lines. Its rendered pixels match the source at the tested 1,300-pixel image size.

Actual searchable result PDF with the same visible checklist and a new text layer.
Source text
0 characters
Result text
176 characters
Recognition check
6 of 6 sample lines
Read the Text Extracted from This Result
WORKSHOP CHECKLIST
Confirm the room booking.
Print forty copies of the handout.
Test the projector before 09:00.
Save the final attendance sheet.
This is a harmless OCR sample.

Settings used: Page 1; searchable PDF; 200 DPI recognition; single block of text; English; original orientation.

This clean, large-print English sample is not an accuracy benchmark for small print, handwriting, skewed scans or other languages. Proofread names, amounts and dates in your own result.

Checked on . Harmless samples processed in a local Chromium browser with this toolkit. These examples are not a live hosting benchmark.

Settings and Output Checks

SettingWhat to expect
Accepted inputPDF
Processing locationYour browser, on your device
File selectionOne input file
Input sizeUp to 100 MB, subject to device memory and decoded-page limits
Recognition batchUp to 20 selected pages; only those pages appear in the searchable PDF download.

The tool adds a text layer; it does not reconstruct tables as an editable spreadsheet or correct factual content. OCR is not a substitute for reviewing names, account numbers and other precision-sensitive text.

Check Recognition Orientation without Changing the Scan

The rotation setting turns the rendered scan supplied to recognition; the searchable PDF retains the source page’s visible appearance. It is not a replacement for Rotate PDF when you want readers to see the page upright. If an existing text layer contains errors, forcing recognition can add another layer and produce duplicate search results. Plain-text OCR can be a useful checked extraction in that situation. Compare easily confused characters such as 0/O, 1/l and decimal points against the image.

Troubleshooting OCR PDF

SituationWhat to do
A name or number is wrongCompare the recognized text with the image. Correct it in a suitable downstream editor before reuse.
An existing bad OCR layer did not improveChoose Recognize every selected page to add a fresh layer. Existing text is retained, so an existing layer may produce duplicate search results; export plain text if you need a clean extraction.
A long packet is rejectedSplit it into groups of 20 pages or fewer, OCR each group and verify them before merging.

A Practical Example and Final Checks

For a typed inspection sheet, search a known equipment ID in the searchable PDF, then copy that ID and one measurement into a text editor. Compare both with the scan and confirm that selection highlights the intended line. A successful word search alone does not establish correct reading order or reliable numeric extraction.

FAQs

Does OCR Change the Page into Editable Word Content?

The output remains a PDF with an added text layer. Use PDF to Word afterward if a reflowed text document is needed.

Which Languages are Available?

This release includes English, Spanish, Portuguese, German, Indonesian, Russian, French, Italian, Turkish, Arabic, Simplified Chinese, Japanese, Korean, Polish and Hindi recognition models. English is the default; select the language matching the source. OCR recognizes characters but does not translate the document. Other languages need a separately configured OCR workflow.

Why was a Page Skipped?

The default setting keeps pages with an existing text layer. Select Recognize every selected page to run recognition on them too.

Does the Output Include Pages I Did Not Select?

No. The searchable PDF contains only the selected pages. To keep a long document together, recognize batches of up to 20 pages and combine the checked outputs in their original order.

Will Recognition Rotation Turn the Visible PDF Page?

No. It rotates the image supplied to recognition while retaining the source page’s visible appearance in the PDF. Use Rotate PDF first when the saved page itself should appear upright.

Can This Tool Reliably Read Handwriting or Rebuild Tables?

The tool recognizes text in the selected language; it does not promise reliable handwriting recognition or reconstruct table cells. Use a sharp, upright typed scan and verify the wording before taking it into a Word or spreadsheet workflow.

Continue with PDF to Word, PDF to Text, Scan to PDF. For long-term storage requirements, see PDF to PDF/A; a searchable PDF is not automatically PDF/A.