AI product Open source
GuideLLM is an open-source platform for benchmarking and evaluating real-world LLM inference deployments. It simulates end-to-end interactions with OpenAI-compatible and vLLM-native servers, using real or synthetic text and multimodal datasets and configurable execution profiles such as synchronous, concurrent, throughput, constant-rate, Poisson, and sweep workloads. It measures SLO-oriented latency and token statistics, including time to first token, inter-token latency, and end-to-end behavior, then produces console, JSON, CSV, or HTML reports for deployment tuning, capacity planning, and regression tracking. The tool provides both a CLI and API, with multiprocessing, threading, asynchronous execution, and Hugging Face, file, synthetic, or custom data sources.
1 use taken from transcripts — each links to the moment in the video.
Open-source tool used to benchmark a specific model before setting GPU memory utilization.
1 in the library.