Skip to content
Convertto

OCR a PDF — Make a Scan Searchable and Selectable

Recognise the text in a scanned PDF and add it as an invisible, selectable layer over the original pages.

A scanned PDF is a picture of a page: you cannot search it, select from it or copy out of it. OCR reads the pictures and puts the words back. This runs Tesseract in your browser, then draws each recognised word transparently at the exact position it was found — so the document still looks precisely as it did, keeps its original quality and file size, and is now searchable in any reader. Nothing is uploaded, which is the point when the scan is a passport, a contract or a medical record.

Runs in your browser
Method
A transparent text layer over the original pages — nothing is re-rendered
Languages
16 languages plus bilingual combinations
Privacy
Runs entirely in your browser — nothing is uploaded
Cost
Free, unlimited, no sign-up

Frequently asked questions

Does OCR change how my document looks?

Not at all. The original page objects are kept exactly as they were and the recognised words are drawn at zero opacity on top. Tools that rebuild the PDF from images instead will visibly soften a scan and multiply its file size.

Why is it slow?

Recognition is genuinely heavy work and it is happening on your own machine rather than a server farm. Expect a few seconds per page. Start with a page or two to confirm the language is right before running a long document.

The text is wrong in places.

OCR accuracy depends almost entirely on the scan. A clean 300 DPI scan reads near-perfectly; a phone photo at an angle, a fax, or a page with a coffee ring will not. Raising the recognition detail helps with small print, and choosing the correct language helps more than anything else — English models guessing at German produce exactly the nonsense you would expect.

My PDF is not a scan and OCR found nothing useful.

Then it already has a text layer and needs no OCR. The PDF Text Extractor will read it directly, accurately and instantly.

Is my document uploaded?

No. The recognition model is downloaded to your browser and the pages are processed there. The traineddata file is cached, so later runs in the same language start immediately.

How to use the ocr pdf

  1. 1Select your scanned pdf — the file stays on your device and is never uploaded.
  2. 2Enter or paste your pages.
  3. 3Choose the language.
  4. 4Set the recognition detail.
  5. 5Set the discard words below.
  6. 6Choose the produce.
  7. 7Enter or paste your page plan.
  8. 8Press Run, then download the result when it is ready.

Sources & specifications

Embed this tool

Put the working ocr pdf on your own site. It runs in your visitors' browsers exactly as it does here — free, no account, nothing uploaded.

Share this tool

Last updated

More pdf tools