Begin with the source document
Retrieval-augmented generation (RAG) uses retrieved material to provide context for an answer. If important information never makes it into the ingested representation, later retrieval changes may not address the original loss.
Before experimenting with chunk sizes, inspect a few questions about the PDF:
- Do the relevant pages contain extractable text, or mostly raster images?
- Are text columns, sidebars, or captions likely to require layout-aware reading?
- Are important values in tables, including tables that cross page boundaries?
- Do charts, screenshots, forms, or attachments carry task-critical information?
- Does copied text look plausible, or show unusual symbols and duplication?
A preflight report helps locate evidence for these questions. It cannot determine whether a source is true, complete, up to date, or sufficient to answer a particular user's question.
Use the findings to plan extraction
| Observed source condition | Next step to consider |
|---|---|
| Little text and a page-sized image | Inspect the page; consider OCR if its content is primarily words. |
| Columns or complex tables | Review output from a structure-aware extraction approach. |
| Important diagrams or screenshots | Consider visual processing or retain images with useful provenance. |
| Forms, comments, or attachments | Decide explicitly whether these channels belong in the task. |
These are action categories, not provider recommendations. The right choice depends on your source, task, and acceptable tradeoffs.
Check the representation before indexing it
Open your extracted text or structured output next to the source. For a representative subset, verify section coverage, names and numbers, table associations, and page provenance. Include flagged pages and apparently ordinary pages.
Write several questions whose answers you can locate precisely in the source. Include a question about prose, one about a table, and one about visual content if that content matters to your application. Record the source page for each answer.
This is a suggested manual workflow. The current checker does not run your parser or automatically validate its representation.
Only then examine chunking and retrieval
If the needed information exists in the indexed representation, inspect whether retrieval finds the right passages and whether those passages preserve enough context. If the information is absent or misrepresented, return to extraction and source handling.
A good preflight score does not mean “RAG-ready” in a universal sense. It summarizes a defined set of source checks. You still need to evaluate your parser, retrieval setup, and answers against your own task.
Check the source before the pipeline.
Get an evidence-backed PDF profile with no LLM API key.
Check a PDF ↗