AI product Open source · Apache-2.0

NVIDIA Model Optimizer

NVIDIA Model Optimizer (ModelOpt) is an open-source library for optimizing deep-learning and generative-AI models for inference deployment. It accepts Hugging Face, PyTorch, or ONNX models and provides Python APIs for composing post-training quantization, quantization-aware training and distillation, pruning, neural architecture search, speculative decoding, and sparsity techniques, producing optimized or quantized checkpoints.

View repository Visit site Mentioned in 1 video ↓

Overview

The resulting checkpoints can be exported for inference frameworks including TensorRT-LLM, TensorRT, vLLM, and SGLang. The library also integrates with NVIDIA Megatron-Bridge, Megatron-LM, Hugging Face Accelerate, Transformers, and Diffusers for supported training and export workflows. It is distributed as the nvidia-modelopt package on PyPI and can also be installed from the NVIDIA GitHub repository or used through NVIDIA container images.

What NVIDIA Model Optimizer is used for

1 use taken from transcripts — each links to the moment in the video.

  • An open-source library for compressing models and speeding up inference on NVIDIA GPUs. It supports quantization, sparsity pruning, and distillation through Python APIs.

Videos mentioning NVIDIA Model Optimizer

1 in the library.