AI product Open source

LiteParse

LiteParse is an open-source, standalone document parser from LlamaIndex focused on fast, local parsing of PDFs and other documents without proprietary LLM features or cloud dependencies. Its Rust-based pipeline converts Office files and images to PDF when needed, extracts native text with PDFium, optionally applies Tesseract or an HTTP OCR service, merges the results, and reconstructs spatial layout using text positions and bounding boxes. It can output Markdown, JSON, or layout-preserved plain text, generate page screenshots, detect documents that need OCR or heavier processing, and expose classified layout blocks such as headings, paragraphs, lists, tables, code, and figures; the JSON representation can include per-cell table coordinates and other opt-in PDF metadata. LiteParse is distributed through CLI, Rust, Python, Node.js/TypeScript, and browser WebAssembly packages, and is licensed under Apache 2.0.

View repository Visit site Mentioned in 1 video ↓

What LiteParse is used for

1 use taken from transcripts — each links to the moment in the video.

  • Open-source tool from LlamaIndex for structural content extraction from documents, encoding tables, charts and segmented information into formats agents can reason over spatially.

Videos mentioning LiteParse

1 in the library.