AI use cases

Paged attention

Allocate KV-cache memory in fixed-size, non-contiguous pages to reduce fragmentation and improve GPU utilization.

From How KV Cache Speeds Up LLMs for Faster AI Models on GPUs by IBM Technology