AI product Open source

tax-doc-classifier

An open-source tax-document page classifier built with TypeSafe Jev. It extracts text from PDF pages, then uses decision-model requests to classify each page by form and page kind, returning probability distributions, a form identifier, confidence scores, and a gate indicating whether the result meets the configured confidence threshold. Its generated JSON criteria file describes IRS forms using data extracted from IRS form PDFs; a second request handles schedules for selected corporate and foreign forms. The package includes PDF helpers based on Poppler, supports text-layer input, and requires a TypeSafe API key for the Jev backend. It does not train or host a model, is limited to English federal-form classification, requires OCR for scanned pages, and is distributed under Apache-2.0 with separate licensing for IRS-derived data.

View repository Mentioned in 1 video ↓

What tax-doc-classifier is used for

1 use taken from transcripts — each links to the moment in the video.

  • Classifies tax-form pages from PDF text, returning labels and confidence scores so low-confidence results can be routed for review.

Videos mentioning tax-doc-classifier

1 in the library.