AI product Open source
vLLM is an open-source library and inference-serving engine for deploying large language models on GPUs and other hardware accelerators. Originally developed in UC Berkeley's Sky Computing Lab, it is maintained by a broad community and provides an OpenAI-compatible API server, as well as Anthropic Messages API and gRPC support.
Its serving stack manages attention key-value memory with PagedAttention, continuous batching, chunked prefill, prefix caching, streaming generation, and disaggregated prefill, decode, and encode. It also supports speculative decoding, quantization, optimized attention and GEMM/MoE kernels, automatic kernel generation and graph-level transformations, structured outputs, tool calling, multiple decoding algorithms, and distributed inference through tensor, pipeline, data, expert, and context parallelism.
The project integrates with Hugging Face models and supports decoder-only, mixture-of-experts, hybrid state-space, multimodal, embedding, retrieval, reward, and classification models. It can run across NVIDIA, AMD, Intel, and other supported accelerators, as well as x86, ARM, and PowerPC CPUs, and can be installed with uv or pip or built from source.
6 uses taken from transcripts — each links to the moment in the video.
Provides an open-source inference engine used as part of the standard model-support stack.
Runs language models efficiently at production scale on hardware accelerators, supporting batching, paged attention, broad hardware and model compatibility, and OpenAI-compatible endpoints.
Open-source inference engine for turning GPUs and other accelerators into model-serving endpoints, supporting model architectures, optimizing serving, and benchmarking hardware.
Open-source inference engine providing KV cache, paged attention, prefix caching, chunked prefill, and speculative decoding capabilities for scaling LLM inference.
Serves the deployed model on the GPU rack.
An LLM inference engine that loads model weights, manages GPU memory and KV caches, batches requests, and serves an API. The video uses it to demonstrate how quantizing model weights and the KV cache increases concurrent serving capacity.
6 in the library.