AI use cases
Reuse KV-cache memory for requests with shared system prompts and avoid repeated prefill computation.
From How KV Cache Speeds Up LLMs for Faster AI Models on GPUs by IBM Technology