AI product Open source
WARP (Weight-Aware Runtime and Paging), formerly WASTE, is an embeddable inference engine written in dependency-free C by SQLiteAI. It runs large mixture-of-experts models beyond available RAM by keeping the shared model trunk in memory, streaming selected experts from NVMe or other fast storage, and using remaining RAM as a bounded expert cache. Its container layout gives each expert an aligned read; reads overlap computation, while a lookahead router prefetches experts for the next layer without changing the decisions made by the actual router. Experts use 3-bit residual vector quantization, while more sensitive shared weights use 4- or 8-bit storage. The project is designed for local inference of frontier-scale models such as Kimi K3 and GLM-5.3-Flash on consumer hardware.
2 uses taken from transcripts — each links to the moment in the video.
An open-source, dependency-free C inference engine that streams active experts of large mixture-of-experts models from NVMe while keeping the shared trunk in memory, allowing very large models to run locally with limited RAM.
Runs mixture-of-experts models larger than available memory by keeping the shared trunk in RAM and streaming selected experts from NVMe with caching and read-ahead.
2 in the library.