AI use cases
Reuse previously computed key and value matrices during autoregressive token generation.
From How KV Cache Speeds Up LLMs for Faster AI Models on GPUs by IBM Technology