AI product Open source

AirLLM

AirLLM is an open-source Python library for running large language model inference on GPUs with limited memory. It reduces GPU memory use by loading and processing one model layer at a time instead of holding the entire model in GPU memory; for sparse mixture-of-experts models, it streams individual experts as needed. The project is designed to run large models without requiring quantization, distillation, or pruning, and provides a pip-installable package with support for multiple model families, CPU inference, macOS, and optional model compression and quantization features.

View repository Mentioned in 2 videos ↓

What AirLLM is used for

2 uses taken from transcripts — each links to the moment in the video.

  • An open-source Python library for running large language models on small GPUs by keeping only one model layer on the GPU at a time.

  • A Python library for running large language model inference on GPUs with very little memory, without quantization, distillation, or pruning. It loads and processes one layer at a time so memory use depends on layer size rather than total model size.

Videos mentioning AirLLM

2 in the library.