AI product Open source
DwarfStar is a self-contained, model-specific native inference engine for running DeepSeek V4 Flash, GLM 5.2, and, on high-memory systems, DeepSeek V4 PRO locally. It combines model loading, prompt rendering, tool calls, persistent KV state, an HTTP server, and a coding agent rather than serving as a general GGUF runner. The engine supports Metal on Macs, CUDA including multi-GPU systems, and ROCm on Strix Halo hardware; it also provides SSD streaming for machines with insufficient RAM, distributed tensor and pipeline parallelism, and server-side micro-batching for multi-user inference. The repository includes tooling and data for GGUF, imatrix, quality, and speed testing, and acknowledges code and design contributions from llama.cpp and GGML.
1 use taken from transcripts — each links to the moment in the video.
An open-source native inference engine for running DeepSeek V4 Flash and Pro models on local hardware. It supports model loading, prompt rendering, tool calling, persistent KV state, distributed inference, compatible servers, and a coding agent.
1 in the library.