Scanned PDF to Searchable Text: OCR Guide

Scanned PDF to Searchable Text: OCR Guide

You've been there: a "PDF" that's really just photos of pages. You can't search it, can't copy a paragraph, can't even select a word — Ctrl+F finds nothing, and quoting a sentence means retyping it. That file needs OCR (Optical Character Recognition): software that looks at the pictures, recognizes the letters, and layers real text back into the document. Here's how to do it free.

How to tell if your PDF needs OCR

Quick test: open the PDF and try to select a sentence with your cursor. If the text highlights, it's a real digital PDF — no OCR needed. If your cursor selects the whole page as one big image, or nothing selects at all, it's a scan. Receipts photographed with a phone, old documents run through a flatbed scanner, and "print then scan" contracts are the usual suspects.

Method 1: Google Drive (free, surprisingly good)

Google's OCR is one of the best free options and most people already have an account:

  1. Upload the scanned PDF to Google Drive.
  2. Right-click it and choose Open with > Google Docs.
  3. Docs runs OCR and opens an editable document: the original scan at the top, recognized text below it.
  4. Clean up the text, then File > Download > PDF Document to get a searchable PDF back.

Accuracy on clean scans is excellent — Google handles multiple languages and even mixed layouts well. The trade-off is privacy: your document is uploaded to Google's servers. For bank statements or medical records, consider the browser-based option below instead.

Method 2: Browser-based OCR (private, page by page)

When the document is sensitive, keep it on your device:

  1. Open PDFPax Image to Text (OCR).
  2. Drop in the scanned pages. Recognition runs entirely in your browser — nothing is uploaded.
  3. Copy the extracted text and paste it where you need it: a Word doc, an email, a new PDF via JPG to PDF.

Be upfront about the limits: this extracts the text — it doesn't rebuild a pixel-perfect searchable PDF the way Acrobat does. For the common jobs (pulling an address off a scanned letter, quoting a scanned contract, digitizing notes), extracted text is exactly what you need. And since it never leaves your device, it's the right choice for anything you'd hesitate to upload.

Method 3: Adobe Acrobat (paid, best for big batches)

Acrobat's Scan & OCR > Recognize Text processes entire multi-page scans in one pass and produces a proper searchable PDF — the original page images stay visually identical, with an invisible text layer underneath for searching and selecting. For archiving boxes of old paperwork or processing long reports, nothing free matches the throughput. For a handful of pages, the free methods above are faster than installing anything.

Getting better accuracy: it starts before the OCR

OCR quality is mostly decided by the scan, not the software. A few things that genuinely help:

  • Straighten the page. Crooked scans confuse recognition more than anything else. Most scanner apps auto-deskew — use that feature.
  • Good lighting, flat pages. Phone-scan in daylight, press the page flat, avoid shadows from your hand. Curled book pages are OCR's nemesis.
  • 300 DPI or higher. If your scanner offers a resolution setting, 300 DPI is the sweet spot for text. Below 200, accuracy falls off a cliff.
  • Handwriting: lower your expectations. Printed text converts at 98%+ accuracy on clean scans. Handwriting is a different sport entirely — neat block letters sometimes work, cursive mostly doesn't. No free tool reliably transcribes handwritten pages.
  • Proofread numbers. OCR confuses 0/O, 1/l, 5/S. Always double-check account numbers, dates, and amounts — the one place errors actually cost you.

Once your text is extracted, you might want it in an editable document — our guide to converting PDF to Word picks up where this one leaves off.

FAQ

Will OCR keep my formatting?

Mostly no. OCR recovers the words; columns, tables, and styling are approximated at best. Expect to reformat anything beyond simple paragraphs.

Can OCR handle Urdu, Arabic, or other scripts?

Google's OCR supports many scripts including Urdu and Arabic reasonably well on clean scans. Browser-based tools vary — check the tool's language list before processing a large document.

Is there a page limit on free OCR?

Google Drive's method handles reasonably large files but slows down past a few dozen pages. For very long scans, Acrobat's batch OCR (paid) or splitting the file first with Split PDF and processing in chunks works better.

Have a scan to read? Extract its text free with PDFPax OCR — no sign-up, and your files never leave your device.