AI use cases

KV caching

Reuse previously computed key and value matrices during autoregressive token generation.

From How KV Cache Speeds Up LLMs for Faster AI Models on GPUs by IBM Technology