SSD-streamed large-model inference for memory-limited Macs

Slotstream runs a 125-billion-parameter model by streaming its 104 GB weights from an SSD and exposing an Ollama-compatible API.

From Github AwesomeGitHub Trending Weekly #47: utopia, neo, slotstream, Codewhale, TokensBurned, mono-color-skill at 01:04

Problem: Large models exceed available system memory on many computers.

For: Mac users who want to run large language models without enough RAM to hold the full model.

Products from this video

ABYSSAL cc-prune CDAF (Cached Descriptive Asset Files) Codewhale defragger Editable Visual Design FixAnything genart-skill h3-storyboard-skill HexStellar hqtui Keyword Pro LightNav-0 loadersz markdown-graphs Monocolor Editorial Print Noty Omakade Procedura qwen38-27b-rtx3090 Shrimply Skill Cabinet SkillRadar skin-tokens.cpp slotstream Strata TokensBurned TrustMeBro Utopia vol-rs

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Related ideas