AI product Open source · Apache-2.0
NVIDIA Model Optimizer (ModelOpt) is an open-source library for optimizing deep-learning and generative-AI models for inference deployment. It accepts Hugging Face, PyTorch, or ONNX models and provides Python APIs for composing post-training quantization, quantization-aware training and distillation, pruning, neural architecture search, speculative decoding, and sparsity techniques, producing optimized or quantized checkpoints.
The resulting checkpoints can be exported for inference frameworks including TensorRT-LLM, TensorRT, vLLM, and SGLang. The library also integrates with NVIDIA Megatron-Bridge, Megatron-LM, Hugging Face Accelerate, Transformers, and Diffusers for supported training and export workflows. It is distributed as the nvidia-modelopt package on PyPI and can also be installed from the NVIDIA GitHub repository or used through NVIDIA container images.
1 use taken from transcripts — each links to the moment in the video.
An open-source library for compressing models and speeding up inference on NVIDIA GPUs. It supports quantization, sparsity pruning, and distillation through Python APIs.
1 in the library.