SSD-streamed large-model inference for memory-limited Macs
Slotstream runs a 125-billion-parameter model by streaming its 104 GB weights from an SSD and exposing an Ollama-compatible API.
From Github Awesome — GitHub Trending Weekly #47: utopia, neo, slotstream, Codewhale, TokensBurned, mono-color-skill at 01:04
Problem: Large models exceed available system memory on many computers.
For: Mac users who want to run large language models without enough RAM to hold the full model.
Products from this video
ABYSSAL cc-prune CDAF (Cached Descriptive Asset Files) Codewhale defragger Editable Visual Design FixAnything genart-skill h3-storyboard-skill HexStellar hqtui Keyword Pro LightNav-0 loadersz markdown-graphs Monocolor Editorial Print Noty Omakade Procedura qwen38-27b-rtx3090 Shrimply Skill Cabinet SkillRadar skin-tokens.cpp slotstream Strata TokensBurned TrustMeBro Utopia vol-rs