An open-source, dependency-free C inference engine that streams active experts of large mixture-of-experts models from NVMe while keeping the shared trunk in memory, allowing very large models to run locally with limited RAM.
Runs mixture-of-experts models larger than available memory by keeping the shared trunk in RAM and streaming selected experts from NVMe with caching and read-ahead.