📄 PDF · Updated October 10, 2026 · 8 min read

Make a Scanned PDF Searchable With an Invisible Text Layer

Scan: 0 hits Ctrl+F: rabies 🔎

To make a scanned PDF searchable, run OCR on each page and add the recognized words back as an invisible text layer placed exactly over the printed ones, so the page still looks like the scan but Ctrl+F, copy and screen readers finally work. GrabCast's free OCR PDF does this in your browser in 20 languages and saves a new file ending in -ocr.pdf. Below, a fictional two-page vet visit summary goes from a flat picture to a searchable document, including the spots where OCR still needed a human eye.

🔎 Try OCR PDF now — freeOpen →
OCR PDF result for a scanned two-page vet visit summary: searchable PDF ready, recognized text on the left and word boxes on the page on the right
A scanned two-page record turned into a searchable PDF entirely in the browser.
💡 Why a scan is invisible to search

A phone scanner app or office copier saves each page as a photograph wrapped in a PDF. Your eyes see words, but the file contains only pixels, so searching for a patient name, a policy number or the word rabies returns nothing, and selecting text grabs the whole image. Cloud drives, email archives and desktop search tools skip those files too. Adding a hidden text layer fixes every one of those problems without changing how the page looks when printed or shared.

What a searchable scanned PDF actually contains

After OCR the page has two layers. The bottom one is the original scan, untouched. On top sits text drawn in invisible render mode, positioned and sized so each word lines up with its printed twin. Viewers ignore it when painting the page but use it for search, selection, copy and accessibility.

That is different from extracting text into a separate file. Extraction gives you words without the page; a searchable PDF keeps the original page as evidence and adds the words underneath, which is what archives, HR records and insurance claims usually need.

How OCR PDF makes a scanned PDF searchable in the browser

The tool opens the file with pdf.js, reads each page with the open-source Tesseract.js engine and writes the new layer with pdf-lib. The engine and the language data download the first time you use them; the document itself stays on your device.

Our two-page sample, scanned at about 200 DPI with a slight tilt and some paper noise, finished in a few seconds: 2 pages made searchable, 136 words, 88% average confidence, 1.1 MB.

Where OCR still needs a human check

An average confidence of 88% sounds high, and for headings and paragraphs it was: every sentence of the exam notes came through cleanly. The preview, which outlines each recognized word in purple, showed two weaknesses worth knowing.

Scroll the preview before you rely on the file. If a critical block has no purple boxes, a rescan at 300 DPI with more contrast, or cropping the table into its own page, often fixes it. For documents where one wrong digit matters, such as dosages or account numbers, keep treating the scan image as the source of truth.

Using the searchable file afterwards

Download the searchable PDF and try it immediately: open it, press Ctrl+F or Cmd+F and search for a word you can see on page 2.

There is no fixed page limit, but everything runs in the browser's memory, so a 300-page archive box is best done in chunks with the Pages field, such as 1-50, then 51-100.

Once searchable, the Chat with PDF tool lets you ask questions about the content.

Step-by-step

1234
1Set the language — Keep English or add the languages printed in the document, up to three, and leave Pages empty to read every page.
OCR PDF settings: English chosen as the document language, an empty Pages field meaning all pages, and the PDF drop zone below
Check the document language (English is preselected) and leave Pages empty to read every page.
2Drop the scan — Add the PDF; the engine loads once, then each page is read on your device with a progress bar and time estimate.
OCR finished for vet-visit-scan-ocr.pdf: 2 pages made searchable, 136 words, 88% avg. confidence, 1.1 MB, with a Download searchable PDF button
When both pages are read, the summary shows words found, average confidence and the new file size.
3Check the preview — Look at the purple word boxes and the recognized text, and note any tables or numbers that need proofreading.
Page 1 of the scanned Riverbend Animal Clinic summary with headings and paragraphs outlined in purple, while the typed table cells have no boxes
Purple boxes mark each word placed in the invisible text layer; unmarked areas, like these table cells, were not read.
4Download and test — Click Download searchable PDF, open the -ocr.pdf file and search for a word from the last page.
Recognized text panel with Riverbend Animal Clinic, Reason for visit and Exam findings, where 26.4 lb came out as 26.4 1b
Copy the text or save it as .txt, and proofread numbers: here lb was read as 1b.

Common mistakes to avoid

⚠️Leaving the language on English for a document printed in another language, which produces low confidence and garbled words.
⚠️Assuming every block was read; tables, stamps and faint areas can be skipped, so check the word boxes in the preview.
⚠️Running OCR on a digitally signed original and then discovering the signature no longer validates.
⚠️Trusting OCR digits blindly for dosages, amounts or account numbers instead of checking them against the image.

Pro tips

✓Scan at 300 DPI in grayscale when you can; it gives the engine far more to work with than 150 DPI color.
✓Flatten curled pages and avoid shadows from the phone; even lighting improves accuracy more than any setting.
✓Split very long files with the Pages field so a browser tab on an older laptop never runs out of memory.
✓Search the output for a word from every page type, a heading, a paragraph and a table, before archiving it.
✓Keep the original scan too, and name the new file clearly; the tool adds -ocr to the name for you.

Frequently asked questions

Does making a PDF searchable change how it looks?

No. The original page images are kept as they are, and the recognized words are added as invisible text on top, so printing and viewing look the same.

Is the scanned PDF uploaded?

No. pdf.js, Tesseract.js and pdf-lib run in your browser. Only the OCR engine and language data are downloaded the first time; your document stays on your device.

How accurate is the text layer?

Clean typed scans at 200 to 300 DPI are usually close to word-perfect for paragraphs. In our test, table cells were skipped and one l became a 1, so proofread anything critical.

Can I make only some pages searchable?

Yes. Type a range such as 1-3, 5, 8-end in the Pages field. Pages outside the range are left as they are.

What happens to pages that already have text?

Pages with 30 or more characters of selectable text are skipped so the words are not doubled. Tick Also OCR pages that already have selectable text to include them.

📌 Bottom line

A searchable scanned PDF keeps the page exactly as scanned and hides the recognized words underneath it, so search, copy and screen readers work. Pick the right language, let the tool read the pages locally, check the purple word boxes for skipped tables, and test the -ocr file with a quick search.

Open OCR PDF →

Related guides

Browse more: all PDF guides · OCR PDF