How to Make a Scanned PDF Searchable
You have a PDF of a scanned document. It opens fine, it looks fine, and it is completely useless to search. Pressing Ctrl+F finds nothing, you cannot select a sentence to quote, and copying a paragraph is impossible. The file is a stack of photographs wearing a document's clothing.
The fix is a searchable PDF: the same page image you already have, with a layer of recognised text laid invisibly on top. It looks identical and behaves like a document.
Two kinds of PDF that look the same
A PDF made from a word processor contains the actual characters. The letters are stored as text with a font, which is why search and selection work and why the file is usually small.
A PDF made from a scanner or a photograph contains a picture of a page. There are no characters in the file at all, only pixels. Every visual cue tells you it is a document, and none of the machinery behind it agrees.
You can tell them apart in two seconds. Open the file and try to select a line of text with your cursor. If you get a text selection, the characters are there. If you get a rectangular marquee over an image, they are not.
What "searchable" actually means
The trick is layering. The page image stays exactly as it is and remains what you see. Behind it — or strictly, on top of it but rendered invisibly — sits the recognised text, with each word positioned over the place it appears in the image.
Because the text is really there in the file, search finds it, selection selects it, and indexing tools can read it. Because it is rendered invisibly, you never see it and the page looks untouched. When you drag your cursor across a line, you are selecting the invisible text while looking at the picture, and the two line up.
This is why a searchable PDF is usually the right output for a scanned document, rather than extracting the text to a separate file. You keep the original appearance — the letterhead, the signature, the layout, the stamp — and gain the ability to search it. An extracted text file gains searchability by throwing the document away.
Making one
The steps are the same whatever tool you use.
- Get a good capture. Everything downstream depends on this; a blurry scan produces a searchable PDF whose search does not work, which is arguably worse than an honest image because you will trust it.
- Run recognition, and check the text it produced. This is the step people skip. If the recognised text is wrong, the invisible layer is wrong, and nothing about the file will tell you.
- Export as a searchable PDF, which writes the image and the positioned text layer into one file.
- Verify by opening it and searching for a word you know is on the page.
In Textquill this is the searchable PDF option in the Download menu, available after scanning an image. The file is assembled on your own machine from the words just recognised, so a scanned contract or medical letter is never uploaded to build it.
What to check before you trust it
A searchable PDF is only as good as the recognition behind it, and its failure mode is quiet. The document still looks perfect, so there is nothing to alert you that the text layer is unreliable.
Three checks take under a minute. Search for a word you can see in the middle of the page, not the heading, since headings are large and read most reliably. Select a paragraph and paste it somewhere to see what you actually get. If the document contains numbers that matter — an invoice total, a reference, a date — check those specifically, because digits have fewer contextual clues than words and are where errors concentrate.
Limits worth knowing
The invisible text layer relies on a font to position the characters, and the practical consequence is that scripts outside the Latin alphabet may not be included in the layer. A document in Chinese, Japanese, Hindi or Arabic can be recognised perfectly and still produce a PDF whose search does not find those words. The image is unaffected and the recognised text is still available through other exports, but do not assume a searchable PDF is searchable in every language.
Handwriting is a separate limit. Printed type is reliable, neat handwriting is inconsistent, and cursive is largely beyond current recognition. A searchable PDF of a handwritten letter will look right and search badly.
Finally, the text layer is a snapshot. If you later edit the underlying image, the invisible text does not follow.
Why it is worth doing
The value shows up later. A folder of scanned invoices you can search by supplier name is a different thing from a folder of pictures named by date. Scanned research papers become quotable without retyping. Archived correspondence becomes findable by the phrase you half-remember rather than by the filename you never chose carefully.
It also makes documents accessible. A screen reader can read a searchable PDF aloud and cannot read an image-only one at all, which for some readers is the difference between a usable document and a blank one.
For anything you intend to keep, the few seconds spent producing a searchable PDF instead of a plain scan is repaid the first time you need to find something in it.
Try it yourself
Textquill extracts text from any image right in your browser — private, offline, and on your device.
Add Textquill to Chrome