← All guides

TEXT LAYERS / 5 MIN READ

Does my PDF
need OCR?

Start by checking whether the page contains usable text, a picture of text, or both.

A page can look like text without containing text

A scanned letter may be a single image inside a PDF. You can read the words on screen, but a text extractor may have no characters to return. Optical character recognition (OCR) attempts to recognize those words from the image.

A digital PDF can also contain a screenshot, photograph, or diagram next to selectable text. The existence of some extractable text does not establish that every important part of the page is covered.

A quick manual check

  1. Open a representative page in your PDF viewer.
  2. Try selecting a sentence and copying it into a plain-text editor.
  3. Check whether the pasted words match the visible sentence.
  4. Repeat on pages with different layouts or sources.

This is a useful spot check, not a complete assessment. A document can mix digital pages with scanned inserts, and selectable text can itself be corrupted.

What the checker can tell you

CheckDocs flags a likely scanned page when raster images cover at least 80% of the page and fewer than 40 non-whitespace text characters can be extracted. The evidence includes the page number, image coverage, and character count.

MEASURED EXAMPLE / PAGE 2
Extractable text
0 characters
Image coverage
100%
Interpretation
Likely scan-like content

These are the measurements in our generated mixed-document example. Its first page contains digital text; its second page contains only a raster image.

A full-page photograph can produce the same signal. The checker cannot know from image geometry alone whether OCR is appropriate. Inspect the page to decide whether it contains words you need to recover.

A missing-text finding can also come from vector content. Text converted to outlines can look readable while yielding few extractable characters. That case is not necessarily a raster scan.

What if the PDF already has OCR?

Some scanned PDFs contain a hidden text layer produced by earlier OCR. They may no longer trigger the scan heuristic. That is useful, but it does not establish that names, numbers, punctuation, or reading order were recognized correctly.

Copy a few important values and compare them visually. Review suspicious-text or overlapping-text findings if present. Avoid treating a high preflight score as an OCR accuracy measurement.

Choose the smallest useful next step

  • Text-heavy scan: consider OCR, then check the resulting text against important source passages.
  • Chart, screenshot, or diagram: consider visual processing or preserve the page image alongside extracted text.
  • Usable digital text: review layout and tables before assuming another OCR pass will help.

CheckDocs recommends these classes of action. It does not run OCR or alter the file.

Find scan-like pages in your PDF.

Get page-level evidence before deciding what needs OCR.

Check a PDF ↗