-
Provides an open-source inference engine used as part of the standard model-support stack.
-
Runs language models efficiently at production scale on hardware accelerators, supporting batching, paged attention, broad hardware and model compatibility, and OpenAI-compatible endpoints.
-
Open-source inference engine for turning GPUs and other accelerators into model-serving endpoints, supporting model architectures, optimizing serving, and benchmarking hardware.
-
Open-source inference engine providing KV cache, paged attention, prefix caching, chunked prefill, and speculative decoding capabilities for scaling LLM inference.