AI use cases

Generate draft tokens for speculative decoding

Allow the larger model to verify multiple proposed tokens in one forward pass and increase decode speed

From The Inference Frontier: 10x Faster Models to Self-Optimizing AI — Philip Kiely & Ali Taha, Baseten by Latent Space