Quantization recipe for fitting a 27-billion-parameter model on a 24 GB GPU
qwen38-27b-rtx3090 re-quantizes two untied 2.5 GB embedding matrices to int8, freeing 2.6 GB for larger batch sizes.
From Github Awesome — GitHub Trending Weekly #47: utopia, neo, slotstream, Codewhale, TokensBurned, mono-color-skill at 13:46
Problem: A 27-billion-parameter model may load on a 24 GB card but leave insufficient memory for useful batched inference.
For: Local-inference users and developers trying to run large language models on single gaming GPUs.
Products from this video
ABYSSAL cc-prune CDAF (Cached Descriptive Asset Files) Codewhale defragger Editable Visual Design FixAnything genart-skill h3-storyboard-skill HexStellar hqtui Keyword Pro LightNav-0 loadersz markdown-graphs Monocolor Editorial Print Noty Omakade Procedura qwen38-27b-rtx3090 Shrimply Skill Cabinet SkillRadar skin-tokens.cpp slotstream Strata TokensBurned TrustMeBro Utopia vol-rs
Examples
- Qwen 3.8 27B: the model being fitted to the single 24 GB card.