1. Placement is not a reading sequence
A PDF describes how content is placed on a page. Text laid out in two columns may be returned in an order that differs from the sequence a reader expects. A sidebar or caption can land in the middle of a paragraph.
Visible layout: a left-column discussion of revenue and a right-column discussion of expenses.
Possible extraction: one revenue line, one expense line, then another revenue line.
Every word may be present while the relationship between sentences becomes unclear.
Preflight looks for substantial text regions separated by a central gutter and sharing vertical space. A finding points to a possible reading-order risk; it does not reconstruct the correct sequence. Sidebars can trigger the rule, and asymmetric layouts may be missed.
2. A table carries relationships, not just values
Consider a table with columns for a year, a product, and a total. Flattened text may preserve “2025,” “Product A,” and “120” but lose which header belongs to which value.
Page breaks create another uncertainty: does the table at the top of the next page continue the earlier one? Repeated column positions near adjacent page boundaries are a useful clue, not proof.
Preflight reports detected table geometry and possible cross-page continuations. Inspect the affected pages and consider a structure-aware parser. Borderless tables, merged cells, and tables embedded in scans can escape these geometric checks.
3. Images can contain the point of the page
A chart can communicate a trend through a line or color legend. A diagram can encode relationships using arrows. Even when the page has a selectable title and caption, ordinary text extraction may omit those relationships.
Image-heavy findings indicate substantial raster-image coverage. They are informational because image presence is normal. A decorative background and a critical screenshot can have similar coverage; a human still needs to judge the role of the visual.
4. Some information lives outside the main text
Review comments, form fields, embedded attachments, and link destinations may matter to a task without appearing in ordinary page text. Cropping can also make rendered content and extracted objects disagree about what is visible.
The checker surfaces presence counts and geometry. It does not execute actions, inspect attachment contents, or infer whether a comment is important. An annotation finding asks you to consider whether that channel belongs in your ingestion workflow.
A practical review method
- Run source preflight and note the affected pages.
- Open a few flagged pages alongside your actual parser output.
- Check column order, table headers and values, captions, and any task-critical annotations.
- Choose an extraction or preservation strategy for the structures that matter.
- Check the result using representative questions with answers you can verify in the source.
This browser checker identifies plausible source risks. It does not compare parser output, establish semantic completeness, or select a best parser. Local representation comparison is available through the Python library and CLI, with lexical checks rather than a semantic completeness guarantee.
Start with the pages worth reviewing.
See source signals for text, tables, visuals, and page geometry.
Check a PDF ↗