AI product Open source
lumabri is a pure-C system for running large mixture-of-experts models across a swarm of peer machines using the colibri engine. A machine hosting a model serves it to other machines; model bytes needed during inference are fetched from peers on first use and retained in a shared content-addressed store. With Segment available, peers retain assigned layer ranges and their state while clients receive the tokenizer, embeddings, final transform or head, and conversation state; the system can instead fall back to the expert/CAS engine or local execution when a complete Segment route is unavailable. It supports CPU- and SSD-first participation as well as GPUs, and provides serving and chat commands, a terminal UI, configurable disk and compute donation, machine and health diagnostics, swarm monitoring, and low-priority peer-to-peer or signed relay transfers. The video also describes routing of missing model blocks or expert activations between peers, with hashes, signatures, and optional encryption used for protection.
1 use taken from transcripts — each links to the moment in the video.
Distributes mixture-of-experts inference across ordinary machines instead of requiring each participant to hold the entire model. It routes missing model blocks or expert activations between peers and uses hashes, signatures, and optional encryption to protect data.
1 in the library.