AI product Open source
kimi-k3-mlx is an MLX port for Apple Silicon of Moonshot’s Kimi K3 native-multimodal mixture-of-experts model. It provides a stock `mlx-lm`-compatible text model definition, a 3D video-capable MoonViT vision tower with a patch-merger projector, and an `mlx-vlm`-shaped wrapper that joins the language and vision components. The model has a 1-million-token context and uses routed and shared experts, Kimi Delta Attention, gated multi-head latent attention, Attention Residuals, latent-dimension experts, and the SiTU-GLU activation; routed experts are distributed in MXFP4 while other weights use bf16.
The multimodal processor expands each image placeholder into a feature block rather than performing a same-length one-token-per-image scatter. The repository includes a streaming converter because standard `mlx-lm` conversion would materialize the multi-terabyte bf16 model, along with REAP expert-pruning scripts and expert-overlap analysis using Chinese, English, code, or mixed calibration data. It also contains vision-parity tests against Moonshot’s reference implementation.
The repository states that the full model cannot run on a single Mac: even its smallest deployment tier requires hundreds of gigabytes of memory and exceeds the capacity of the largest Apple Silicon machine.
2 uses taken from transcripts — each links to the moment in the video.
Ports MoonShot's large model to Apple Silicon using MLX. The project examines how pruning experts with Chinese, English, code, or mixed calibration data affects the model's capabilities.
An Apple Silicon MLX port of Moonshot's Kimi K3 multimodal mixture-of-experts model. It includes text and vision components, streaming quantization, and expert pruning for language or domain-specific use.
2 in the library.