AI product Open source

DocJev

DocJev is an open-source Python library, CLI, and local app for classifying and splitting PDF, DOCX, and PPTX files with user-written natural-language rules. LiteParse extracts page text locally, then the hosted Jev decision engine assigns document categories or identifies boundaries between component documents; optional LlamaParse tiers provide cloud OCR for difficult inputs. Classification returns a category, probabilities, and review flags. Splitting returns ordered categories and contiguous page ranges, with optional export of each segment as a PDF and review reasons for uncertain boundaries. DOCX and PPTX processing requires LibreOffice, and inference is not offline because Jev is a hosted service. The project is distributed under the Apache-2.0 license.

View repository Mentioned in 1 video ↓

What DocJev is used for

1 use taken from transcripts — each links to the moment in the video.

  • Classifies PDFs, Word files, and slide decks using user-written rules, then splits combined document packets into page ranges. It flags uncertain boundaries and can export the separated documents as PDFs.

Videos mentioning DocJev

1 in the library.