AI product Open source

kimi-k3-mlx

kimi-k3-mlx is an MLX port for Apple Silicon of Moonshot’s Kimi K3 native-multimodal mixture-of-experts model. It provides a stock `mlx-lm`-compatible text model definition, a 3D video-capable MoonViT vision tower with a patch-merger projector, and an `mlx-vlm`-shaped wrapper that joins the language and vision components. The model has a 1-million-token context and uses routed and shared experts, Kimi Delta Attention, gated multi-head latent attention, Attention Residuals, latent-dimension experts, and the SiTU-GLU activation; routed experts are distributed in MXFP4 while other weights use bf16.

View repository Mentioned in 2 videos ↓

Overview

The multimodal processor expands each image placeholder into a feature block rather than performing a same-length one-token-per-image scatter. The repository includes a streaming converter because standard `mlx-lm` conversion would materialize the multi-terabyte bf16 model, along with REAP expert-pruning scripts and expert-overlap analysis using Chinese, English, code, or mixed calibration data. It also contains vision-parity tests against Moonshot’s reference implementation.

The repository states that the full model cannot run on a single Mac: even its smallest deployment tier requires hundreds of gigabytes of memory and exceeds the capacity of the largest Apple Silicon machine.

What kimi-k3-mlx is used for

2 uses taken from transcripts — each links to the moment in the video.

  • Ports MoonShot's large model to Apple Silicon using MLX. The project examines how pruning experts with Chinese, English, code, or mixed calibration data affects the model's capabilities.

  • An Apple Silicon MLX port of Moonshot's Kimi K3 multimodal mixture-of-experts model. It includes text and vision components, streaming quantization, and expert pruning for language or domain-specific use.

Videos mentioning kimi-k3-mlx

2 in the library.