AI product Open source
Kimi K3 in C is a portable C99 inference engine for running the 2.78-trillion-parameter Kimi K3 base model on CPUs without BLAS, an inference framework, or a GPU. It reads the 1.56 TB checkpoint from disk, keeps a selected depth of the 93-layer dense trunk in memory, and streams the remainder during generation; its routed experts are packed in 4-bit form, multiplied directly from that representation, and managed with an LRU cache. The implementation includes KDA attention with a non-growing memory, MLA attention with a latent representation, and selection of 16 experts from 896. The repository provides a command-line interface, weightless and full-checkpoint validation gates, and measurement data; it reports operation from 8 GB of RAM upward, with additional memory changing speed rather than the generated output. The base model has no chat template, so prompts produce continuations rather than chat replies. Reference requirements include an AVX2-and-FMA CPU, about 1.7 TB of free storage, and a C compiler; Linux/x86-64 is the reference platform, while macOS/arm64 and Windows/x86-64 are also documented as building and passing the tests.
2 uses taken from transcripts — each links to the moment in the video.
A C99 inference engine that runs Moonshot's 2.78-trillion-parameter Kimi K3 base model on an x86-64 CPU by streaming model data from disk and caching routed experts without a GPU, BLAS, or inference framework.
Runs Moonshot's 2.78-trillion-parameter model through a C99 inference engine without a GPU or machine-learning framework. It streams model data from disk and caches routed experts within a RAM budget.
2 in the library.