AI product Open source

mini-AGI

mini-AGI is an experimental continual-learning, byte-level language model that trains from scratch on a CUDA-capable GPU with at least 8 GB of VRAM. It reads 256 byte values directly rather than using a conventional tokenizer, processes data one chunk at a time, and uses the same forward path for training and generation. Its architecture combines dense prelude blocks with a recurrent block applied up to 24 times, adaptive PonderNet-style halting, and per-application top-8 routing through a dynamically growing and pruning expert pool. Expert weights and Adam moments are stored as files on disk; a working set is paged through RAM and VRAM as needed, allowing the pool size to exceed available VRAM. The project includes Python commands for building corpora, reading files, streaming training, serving a local web interface, and inspecting or replicating continual-learning experiments. The repository describes the current model as toy-level rather than frontier-capable, and its weights are generated in the local weights directory rather than distributed with the source.

View repository Mentioned in 2 videos ↓

What mini-AGI is used for

2 uses taken from transcripts — each links to the moment in the video.

  • A small language model trained from scratch on an 8 GB GPU. It reads raw bytes without a tokenizer, pages experts from disk, and continually learns from new data.

  • An open-source continual-learning language model that trains from scratch on a single 8GB laptop GPU and continues learning from what it reads. It stores weights on disk, pages them to the GPU as needed, grows its capacity, and prunes unused knowledge.

Videos mentioning mini-AGI

2 in the library.