AI product Open source

SGLang

SGLang is an open-source serving framework and inference engine for large language models and multimodal models, developed by the SGLang project. It provides an alternative serving stack for deployed models, with runtime and kernel optimizations for inference workloads.

View repository Visit site Mentioned in 2 videos ↓

Overview

Its documented capabilities include RadixAttention, a zero-overhead batch scheduler, cache-aware load balancing, structured-output decoding, speculative decoding, disaggregated prefill and decode, model parallelism, and support for GPU and TPU backends. The project also publishes integrations and optimizations for current open models and multimodal, image, and video-generation workloads.

What SGLang is used for

2 uses taken from transcripts — each links to the moment in the video.

  • Provides an open-source inference engine whose runtime behavior and kernel issues are compared with other serving engines.

  • Provides an alternative inference-serving stack for the deployed model.

Videos mentioning SGLang

2 in the library.