Tool Open source
PDF Inspector is an open-source Rust library by Firecrawl for classifying PDF documents and extracting position-aware text without OCR. It samples content streams to classify documents as text-based, scanned, image-based, or mixed, returns confidence scores and per-page OCR-routing information, and detects broken font encodings. Its extractor uses font information and X/Y coordinates to determine reading order, including multi-column and right-to-left layouts, and converts content to Markdown with headings, lists, code blocks, tables, formatting, links, and page breaks. Native Rust and CLI consumers can opt into selective OCR for pages that need it; the Python and Node.js bindings can run the same workflow, while browser WebAssembly provides local parsing without a server round trip. The project also supports CID font decoding and shares a single parsed document between classification and extraction.
1 use taken from transcripts — each links to the moment in the video.
A fast open-source Rust library that classifies PDFs and extracts position-aware text without OCR. It identifies scanned and mixed documents, reports pages needing OCR, and can produce structured Markdown.
1 in the library.