AI product
NVIDIA Dynamo is a developer toolkit from NVIDIA for managing distributed key–value (KV) context caches used in large‑language‑model inference. It provides KV‑aware routing, cache offloading, and prefill/decode disaggregation to move KV cache and related activation/state across GPU clusters, reducing memory pressure and supporting scalable LLM serving. It is intended for integration into GPU‑based inference and serving stacks.
1 use taken from transcripts — each links to the moment in the video.
Acts as a developer toolkit for moving KV cache and other information around a cluster, supporting KV-aware routing, offloading, and prefill/decode disaggregation.
1 in the library.