AI use cases
Use a smaller draft model to propose tokens that a larger model verifies in one forward pass, reducing interactive latency.
From How KV Cache Speeds Up LLMs for Faster AI Models on GPUs by IBM Technology