GPU-free local inference engine for a large language model
A C99 inference engine that runs Moonshot's 2.78-trillion-parameter model without a GPU or ML framework by streaming model data from disk and caching routed experts in RAM.
From Github Awesome — GitHub Trending Monthly #10(2026.08) at 01:55
Problem: Enables local inference without GPU hardware or an ML framework.
For: Developers with sufficient RAM and fast local storage who want to run the model without a GPU.
Products from this video
Amadeus Protocol Node Amagine3D Amazon S3 Anatomy Atelier Anti Slop anydoc ARC-AGI-1 Task Generator arc-task-gen claudish-to-english claudish-to-english code-server Comp AI CRM Comp AI CRM Cumora DeepSeek Harness FrontierAgent FuXi fx Gemma Translator Gemma Translator h3.c IP as Logo IP as Logo Skill JoyAI-Video-Edit Kage Kimi K3 in C macOS MangoDisk Model Context Protocol (MCP) Moli morphicons MyContext npm OpenBot openGym Phone Harness Phone Harness Praxist Python Raspberry Pi 5 React Rust Scroll Craft Sol Advisor SQLite terminal-code Three.js ThreeUI TypeScript unlazy V8 Virtual Mac on iPad walgit Xcode Zig ZSvirt
Examples
- kimi-k3-in-c runs Moonshot's 2.78-trillion-parameter model without a GPU.