AI product Open source
AirLLM is an open-source Python library for running large language model inference on GPUs with limited memory. It reduces GPU memory use by loading and processing one model layer at a time instead of holding the entire model in GPU memory; for sparse mixture-of-experts models, it streams individual experts as needed. The project is designed to run large models without requiring quantization, distillation, or pruning, and provides a pip-installable package with support for multiple model families, CPU inference, macOS, and optional model compression and quantization features.
2 uses taken from transcripts — each links to the moment in the video.
An open-source Python library for running large language models on small GPUs by keeping only one model layer on the GPU at a time.
A Python library for running large language model inference on GPUs with very little memory, without quantization, distillation, or pruning. It loads and processes one layer at a time so memory use depends on layer size rather than total model size.
2 in the library.