AI product Open source
An open-source tax-document page classifier built with TypeSafe Jev. It extracts text from PDF pages, then uses decision-model requests to classify each page by form and page kind, returning probability distributions, a form identifier, confidence scores, and a gate indicating whether the result meets the configured confidence threshold. Its generated JSON criteria file describes IRS forms using data extracted from IRS form PDFs; a second request handles schedules for selected corporate and foreign forms. The package includes PDF helpers based on Poppler, supports text-layer input, and requires a TypeSafe API key for the Jev backend. It does not train or host a model, is limited to English federal-form classification, requires OCR for scanned pages, and is distributed under Apache-2.0 with separate licensing for IRS-derived data.
1 use taken from transcripts — each links to the moment in the video.
Classifies tax-form pages from PDF text, returning labels and confidence scores so low-confidence results can be routed for review.
1 in the library.