AI product Open source
Docling is an open-source document-processing toolkit from the Docling project for preparing files for generative-AI applications. It parses formats including PDF, DOCX, PPTX, XLSX, HTML, EPUB, images, plain text, audio, and video, and represents the results in a unified DoclingDocument structure.
For PDFs, it analyzes page layout, reading order, headings, tables, code, formulas, images, and scanned content through OCR. Audio and video inputs can be processed with automatic speech-recognition models; video parsing produces an ASR transcript and representative keyframes. Parsed content can be exported as Markdown, HTML, WebVTT, DocLang, DocTags, or lossless JSON.
Docling supports local and air-gapped execution, a command-line interface, Python usage, an API server named docling-serve, and an MCP server for connecting agents. The project also provides integrations for LangChain, LlamaIndex, CrewAI, and Haystack. It is installable with pip and runs on macOS, Linux, and Windows on x86_64 and arm64 systems.
2 uses taken from transcripts — each links to the moment in the video.
Parses video and audio inputs, wraps Whisper and FFmpeg, extracts transcript timestamps, and exports the result to Markdown.
Converts a PDF into a structured document tree, reconstructing sections, headings, reading order, and tables.
2 in the library.