AI product

NVIDIA Dynamo

NVIDIA Dynamo is a developer toolkit from NVIDIA for managing distributed key–value (KV) context caches used in large‑language‑model inference. It provides KV‑aware routing, cache offloading, and prefill/decode disaggregation to move KV cache and related activation/state across GPU clusters, reducing memory pressure and supporting scalable LLM serving. It is intended for integration into GPU‑based inference and serving stacks.

Visit site Mentioned in 1 video ↓

What NVIDIA Dynamo is used for

1 use taken from transcripts — each links to the moment in the video.

  • Acts as a developer toolkit for moving KV cache and other information around a cluster, supporting KV-aware routing, offloading, and prefill/decode disaggregation.

Videos mentioning NVIDIA Dynamo

1 in the library.