AI product Open source
SGLang is an open-source serving framework and inference engine for large language models and multimodal models, developed by the SGLang project. It provides an alternative serving stack for deployed models, with runtime and kernel optimizations for inference workloads.
Its documented capabilities include RadixAttention, a zero-overhead batch scheduler, cache-aware load balancing, structured-output decoding, speculative decoding, disaggregated prefill and decode, model parallelism, and support for GPU and TPU backends. The project also publishes integrations and optimizations for current open models and multimodal, image, and video-generation workloads.
2 uses taken from transcripts — each links to the moment in the video.
Provides an open-source inference engine whose runtime behavior and kernel issues are compared with other serving engines.
Provides an alternative inference-serving stack for the deployed model.
2 in the library.