AI product Open source

Colibrì

Colibrì is a pure-C inference engine and open research platform for running large mixture-of-experts models on consumer and heterogeneous hardware without engine dependencies. It treats VRAM, system RAM, and storage as a unified multitier inference hierarchy, streaming expert weights from disk and coordinating placement, storage I/O, scheduling, kernels, speculation, and CPU/GPU overlap. The project provides chat, server, and web front ends, with support for CPU, CUDA, Metal, and dual-SSD streaming; its design rules prohibit silently changing model precision or router semantics when fast memory is insufficient.

View repository Visit site Mentioned in 2 videos ↓

What Colibrì is used for

2 uses taken from transcripts — each links to the moment in the video.

  • Runs large mixture-of-experts models by treating VRAM, system memory, and NVMe as a unified weight hierarchy. It supports CPU, CUDA, Metal, and dual-SSD streaming.

  • Runs very large mixture-of-experts models by keeping the dense portion in RAM and streaming routed experts from disk as needed. Written in C, it learns and caches frequently used experts without changing model precision.

Videos mentioning Colibrì

2 in the library.