How OCR makes a scanned PDF searchable
A scan may look like text while containing only pixels. OCR analyzes those pixels and produces characters you can search and copy.
Visible words are not always text
A scanner or phone camera usually creates an image. When that image is placed in a PDF, the page looks normal but search, copy and accessibility tools cannot understand the words. OCR—optical character recognition—estimates which characters appear in the image.
A quick test is to drag across a sentence. If individual words cannot be selected, the page probably needs OCR. Some PDFs mix real text and scanned pages, so test more than the first page.
Input quality controls recognition quality
Straight pages, even lighting and dark text on a clean background produce the best results. Motion blur, shadows, handwriting, decorative fonts and folded paper introduce ambiguity. Rotating a sideways page before OCR is one of the simplest improvements.
Resolution also matters, but larger is not always better. A clear scan around normal document resolution is more useful than a huge, blurred photograph. Crop away unrelated backgrounds when possible.
Treat OCR as a draft
Recognition can confuse similar characters such as zero and O, one and l, or punctuation marks. Names, account numbers and legal clauses deserve manual comparison with the image. OCR output should support discovery and editing, not silently replace the source of record.
The local OCR tool recognizes English and creates a clean text PDF. It runs in the browser, so a long scan can take time and use significant memory. Process smaller sections if a mobile device struggles.