What the checker looks for
The analyzer opens the source PDF once and shares extracted text and layout information across its detectors. It does not run OCR, convert the document, or call a parser vendor.
| Check | Signal | Important limitation |
|---|---|---|
| Weak text layer | Fewer than 40 non-whitespace characters plus image or vector content. | Decorative pages can trigger; blank pages are excluded. |
| Likely scan | Fewer than 40 characters with at least 80% raster-image coverage. | A photograph can look scan-like. Existing OCR can mask a scan. |
| Rotation | Nonzero PDF page rotation. | Intentional rotation is common. Landscape dimensions alone do not trigger. |
| Suspicious text | High unusual-character or symbol ratios. | Formulas and legitimate symbols may trigger. |
| Duplicate text | Repeated text lines with substantial geometric overlap. | Intentional overprinting may trigger; near-identical OCR may be missed. |
| Visual content | At least 40% image coverage. | Informational. Background images count; vector-only charts are not covered by this check. |
| Columns | Substantial text on both sides of a central gutter with overlapping vertical ranges. | Sidebars can trigger. Asymmetric columns may be missed. |
| Tables | Table geometry, plus matching columns near adjacent page edges. | Borderless and scanned tables may be missed; continuation is only a candidate. |
| Non-page content | Annotations, widgets, embedded files, and links. | Presence does not establish importance. Contents are not executed. |
| Page geometry | Crop differences, out-of-bounds objects, or unusual page sizes. | Cropping and foldouts may be intentional. |
Every finding comes with a reason
A finding includes affected page numbers, a short observation, why it could matter, and measurements such as text counts, image coverage, or bounding boxes. Page numbers start at 1. Coordinates are PDF points, not screen pixels.
The report includes counts and geometry, not document text snippets, attachment contents, form values, or annotation bodies. Confidence reflects the strength of a detector signal; it is not a probability that an AI model will be correct.
Suggested actions identify a class of next step. For example, “consider OCR” does not mean OCR has been performed or that a specific provider is recommended.
Office document checks
DOCX checks cover tracked revisions, explicit hidden text, text boxes, tables, and package features. PPTX checks cover speaker notes, hidden slides, multiple text-bearing shapes, and package features. XLSX checks cover formulas, missing cached values, stored errors, hidden regions, and merged cells. Evidence uses document-body, slide, or sheet locations.
Office files receive findings without a numeric score. We do not render Office layouts, evaluate formulas, follow external links, or run macros. Package assets may be unused; a finding is a reason to review, not proof of extraction failure. Encrypted, macro-enabled, Strict OOXML, and legacy Office formats are unsupported.
What does the preflight score mean?
The PDF score starts at 100. Each risk category subtracts a severity-weighted penalty that also considers the fraction of affected pages. High, medium, and low weights are 30, 12, and 4. Informational findings carry no penalty. Correlated scan and missing-text findings share a category, so they do not double the penalty.
category penalty = maximum severity weight ×
(0.5 + 0.5 × affected-page fraction)Penalties are summed, rounded once, and the result is limited to 0–100. A small number of important pages can still matter to your task even if the document-level score looks reassuring.
- 90–100: low observed ingestion risk.
- 75–89: some issues worth reviewing.
- 50–74: remediation likely useful.
- 0–49: substantial observed ingestion risk.
A blank PDF can score 100 because no risk fired. Read its empty_or_unreadable classification. If a detector failed, read the warning: the analysis is incomplete regardless of the score.