Scanned PDF to Searchable Text: OCR Guide
In this guide
You've been there: a "PDF" that's really just photos of pages. You can't search it, can't copy a paragraph, can't even select a word — Ctrl+F finds nothing, and quoting a sentence means retyping it. That file needs OCR (Optical Character Recognition): software that looks at the pictures, recognizes the letters, and layers real text back into the document. Here's how to do it free.
How to tell if your PDF needs OCR
Quick test: open the PDF and try to select a sentence with your cursor. If the text highlights, it's a real digital PDF — no OCR needed. If your cursor selects the whole page as one big image, or nothing selects at all, it's a scan. Receipts photographed with a phone, old documents run through a flatbed scanner, and "print then scan" contracts are the usual suspects.
Method 1: Google Drive (free, surprisingly good)
Google's OCR is one of the best free options and most people already have an account:
- Upload the scanned PDF to Google Drive.
- Right-click it and choose Open with > Google Docs.
- Docs runs OCR and opens an editable document: the original scan at the top, recognized text below it.
- Clean up the text, then File > Download > PDF Document to get a searchable PDF back.
Accuracy on clean scans is excellent — Google handles multiple languages and even mixed layouts well. The trade-off is privacy: your document is uploaded to Google's servers. For bank statements or medical records, consider the browser-based option below instead.
Method 2: Browser-based OCR (private, page by page)
When the document is sensitive, keep it on your device:
- Open PDFPax Image to Text (OCR).
- Drop in the scanned pages. Recognition runs entirely in your browser — nothing is uploaded.
- Copy the extracted text and paste it where you need it: a Word doc, an email, a new PDF via JPG to PDF.
Be upfront about the limits: this extracts the text — it doesn't rebuild a pixel-perfect searchable PDF the way Acrobat does. For the common jobs (pulling an address off a scanned letter, quoting a scanned contract, digitizing notes), extracted text is exactly what you need. And since it never leaves your device, it's the right choice for anything you'd hesitate to upload.
Method 3: Adobe Acrobat (paid, best for big batches)
Acrobat's Scan & OCR > Recognize Text processes entire multi-page scans in one pass and produces a proper searchable PDF — the original page images stay visually identical, with an invisible text layer underneath for searching and selecting. For archiving boxes of old paperwork or processing long reports, nothing free matches the throughput. For a handful of pages, the free methods above are faster than installing anything.
Getting better accuracy: it starts before the OCR
OCR quality is mostly decided by the scan, not the software. A few things that genuinely help:
- Straighten the page. Crooked scans confuse recognition more than anything else. Most scanner apps auto-deskew — use that feature.
- Good lighting, flat pages. Phone-scan in daylight, press the page flat, avoid shadows from your hand. Curled book pages are OCR's nemesis.
- 300 DPI or higher. If your scanner offers a resolution setting, 300 DPI is the sweet spot for text. Below 200, accuracy falls off a cliff.
- Handwriting: lower your expectations. Printed text converts at 98%+ accuracy on clean scans. Handwriting is a different sport entirely — neat block letters sometimes work, cursive mostly doesn't. No free tool reliably transcribes handwritten pages.
- Proofread numbers. OCR confuses 0/O, 1/l, 5/S. Always double-check account numbers, dates, and amounts — the one place errors actually cost you.
Once your text is extracted, you might want it in an editable document — our guide to converting PDF to Word picks up where this one leaves off.
FAQ
Will OCR keep my formatting?
Mostly no. OCR recovers the words; columns, tables, and styling are approximated at best. Expect to reformat anything beyond simple paragraphs.
Can OCR handle Urdu, Arabic, or other scripts?
Google's OCR supports many scripts including Urdu and Arabic reasonably well on clean scans. Browser-based tools vary — check the tool's language list before processing a large document.
Is there a page limit on free OCR?
Google Drive's method handles reasonably large files but slows down past a few dozen pages. For very long scans, Acrobat's batch OCR (paid) or splitting the file first with Split PDF and processing in chunks works better.
Have a scan to read? Extract its text free with PDFPax OCR — no sign-up, and your files never leave your device.
Searchable PDF vs. plain text: which output do you need?
OCR gives you a choice, and picking wrong wastes the whole run:
- Searchable PDF — the original page images with an invisible text layer underneath. Looks identical to the scan, but you can search, select, and copy text. Choose this for archives, contracts, and anything where the original appearance matters.
- Plain text (.txt) — just the words, no layout. Choose this when you want to quote, analyze, translate, or feed the content to another tool. It's also much smaller.
- Word document — text plus approximate layout. Choose this when you need to edit the content afterward.
When in doubt, make the searchable PDF — you can always extract plain text from it later, but you can't reconstruct the original look from a .txt file.
Batch OCR workflow for big stacks
Digitizing a filing cabinet? Process matters more than tools:
- Sort before you scan. Group pages by document and remove blanks, duplicates, and sticky notes. Every junk page you OCR is time wasted.
- Scan at 300 DPI, grayscale. Higher DPI barely improves accuracy but explodes file size; color doubles it for no benefit on text pages.
- OCR in batches of 20–50 pages so one bad page doesn't stall the whole job, and spot-check every batch before continuing.
- Verify against the originals for anything important — OCR is ~98–99% accurate on clean scans, which still means a few wrong characters per page. Names, numbers, and dates deserve a human eye.
- Name files for retrieval:
2024-contracts-acme-searchable.pdfbeatsscan003_final2.pdfevery time.
Troubleshooting poor results
- Gibberish output: the scan is too low-resolution or skewed. Re-scan at 300 DPI, straighten the page, and try again — no OCR engine rescues a bad scan.
- Right language, wrong characters: check the OCR language setting. Running English OCR on Urdu, Arabic, or French text produces nonsense; select every language present in the document.
- Tables become soup: normal. OCR reads text, not table structure. For data extraction, convert the searchable PDF to Excel afterward and rebuild.
- Handwriting barely recognized: expected — handwriting OCR remains unreliable for most engines. Transcribe critical handwritten sections manually.
- File size explodes: OCR adds a text layer but keeps the full page images. Compress the searchable PDF if size matters.
Free vs. paid OCR: when to upgrade
Free OCR (Google Drive, browser tools) handles clean, single-language scans remarkably well. Consider paid OCR when:
- Volume is high: hundreds of pages make per-page free tools impractical; Acrobat or dedicated software batch-processes overnight.
- Accuracy is critical: legal, medical, or financial documents where one wrong digit matters. Paid engines edge out free ones on difficult scans — but still verify.
- Layout must be preserved: multi-column pages, tables, and forms reconstruct better in tools that understand document structure.
- Multiple tricky languages: mixed-language documents and unusual scripts get better language models in paid tools.
For everyone else — students, small offices, personal archives — free OCR plus a careful verification pass is the rational choice. Don't buy software to solve a problem you have twice a year.
Free vs. paid OCR: when to upgrade
Free OCR (Google Drive, browser tools) handles clean, single-language scans remarkably well. Consider paid OCR when:
- Volume is high: hundreds of pages make per-page free tools impractical; Acrobat or dedicated software batch-processes overnight.
- Accuracy is critical: legal, medical, or financial documents where one wrong digit matters. Paid engines edge out free ones on difficult scans — but still verify.
- Layout must be preserved: multi-column pages, tables, and forms reconstruct better in tools that understand document structure.
- Multiple tricky languages: mixed-language documents and unusual scripts get better language models in paid tools.
For everyone else — students, small offices, personal archives — free OCR plus a careful verification pass is the rational choice. Don't buy software to solve a problem you have twice a year.
Frequently asked questions
Can OCR handle handwriting?
Poorly, in most cases. Print-style handwriting sometimes works; cursive rarely does. For handwritten notes, you're better off transcribing the important parts manually.
Does OCR work on photos of documents?
Yes, if the photo is sharp, well-lit, and straight. Phone scans via a scanning app (which crops and de-skews) OCR far better than casual tilted photos.
Key takeaways
- Can't select text in your PDF? It's a scan and needs OCR — conversion alone won't help.
- Scan at 300 DPI grayscale: the single biggest factor in OCR accuracy.
- Set the OCR language to match the document, including every language present.
- Make a searchable PDF by default — you can extract plain text from it later.
- Spot-check names, numbers, and dates: even 99% accuracy leaves errors on every page.