AI product Open source

Gigatoken

Gigatoken is an open-source tokenizer for language-model training data, implemented in Rust and distributed as a Python package. It tokenizes large text corpora at gigabytes per second, supports common tokenizer formats across x86 and ARM hardware, and provides compatibility modes for Hugging Face Tokenizers and Tiktoken. Its native API can read text files directly and encode them with parallel processing, while the compatibility modes are designed as drop-in replacements but incur additional overhead. The repository reports throughput of up to roughly 1,000 times that of Hugging Face tokenizers in its benchmarks.

View repository Mentioned in 1 video ↓

What Gigatoken is used for

1 use taken from transcripts — each links to the moment in the video.

  • An open-source tokenizer for language modeling that processes text at gigabytes per second. It is designed for large training corpora and supports common tokenizers on x86 and ARM, including compatibility with Hugging Face and TikToken code.

Videos mentioning Gigatoken

1 in the library.