AI product Open source

MoonEP

MoonEP is an expert-parallelism communication library from Moonshot AI for mixture-of-experts training and inference. It plans a small number of dynamic redundant experts from the current router outputs, prefetches their weights before expert computation, and reduces the resulting gradients back to the experts' home ranks during the backward pass. Its near-optimal GPU planning kernel, fused permute/unpermute operations, and zero-copy dispatch send tokens directly to expert-grouped positions on remote ranks and return buffer views to computation. The library uses a fixed S × K token buffer per rank, where S is the input-token count and K is the routed top-k, keeping communication and computation shapes static across layers. The repository documents support for NVIDIA GPUs and lists Zhenwu PPU support as under review.

View repository Mentioned in 1 video ↓

What MoonEP is used for

1 use taken from transcripts — each links to the moment in the video.

  • Moonshot AI's expert parallel-communication library for balancing mixture-of-experts training when routing concentrates tokens on a few experts. It plans redundant experts, prefetches weights, and uses zero-copy dispatch across Nvidia GPUs.

Videos mentioning MoonEP

1 in the library.