GPU-free local inference engine for a large language model

A C99 inference engine that runs Moonshot's 2.78-trillion-parameter model without a GPU or ML framework by streaming model data from disk and caching routed experts in RAM.

From Github AwesomeGitHub Trending Monthly #10(2026.08) at 01:55

Problem: Enables local inference without GPU hardware or an ML framework.

For: Developers with sufficient RAM and fast local storage who want to run the model without a GPU.

Products from this video

Amadeus Protocol Node Amagine3D Amazon S3 Anatomy Atelier Anti Slop anydoc ARC-AGI-1 Task Generator arc-task-gen claudish-to-english claudish-to-english code-server Comp AI CRM Comp AI CRM Cumora DeepSeek Harness FrontierAgent FuXi fx Gemma Translator Gemma Translator h3.c IP as Logo IP as Logo Skill JoyAI-Video-Edit Kage Kimi K3 in C macOS MangoDisk Model Context Protocol (MCP) Moli morphicons MyContext npm OpenBot openGym Phone Harness Phone Harness Praxist Python Raspberry Pi 5 React Rust Scroll Craft Sol Advisor SQLite terminal-code Three.js ThreeUI TypeScript unlazy V8 Virtual Mac on iPad walgit Xcode Zig ZSvirt

Examples

🔒 Full analysis locked

Unlock more videos and the full analysis

Buy credits to process more videos. Each run includes the full analysis, not just the summary — and you get access to the locked analysis across the library.

Get credits →