AI product Open source
TurboFieldfare is a model-specific Swift and Metal runtime for running the instruction-tuned Gemma 4 26B-A4B model on Apple Silicon Macs, including systems with 8 GB of RAM. It keeps the model’s shared 1.35 GB core and FP16 key-value cache in memory while streaming the expert weights selected for each token from an SSD, allowing inference with about 2 GB of resident weights instead of loading the approximately 14.3 GB model. The model uses MLX affine 4-bit weights with group size 64, an 8-bit router, 4-bit shared and routed experts, and about 3.88 billion active parameters per token.
The project includes a native macOS app, command-line interface, Swift runtime library with Metal kernels, decode service, streaming model installer and verifier, and experimental loopback OpenAI-compatible Chat Completions server. These products share a repacked `.gturbo` model directory; the installer streams byte ranges from a pinned Hugging Face revision, repacks them without materializing the source checkpoint, and validates the resulting manifest and file hashes. The app and CLI support instruction and completion generation with optional system guidance but do not execute tools; the server accepts function-tool declarations and returns model-produced tool calls for client authorization and execution.
An optional companion vision tower adds image input on M2 or newer Macs, while text-only inference remains available on M1. The package is arm64-only and targets macOS 26, Metal 4, and Swift 6.2 or newer, with approximately 14.3 GB of storage for the text model plus about 1.1 GB for the image pack.
3 uses taken from transcripts — each links to the moment in the video.
Runs Gemma 4's mixture-of-experts model on an 8 GB Apple Silicon Mac without loading all weights into memory. Its Swift and Metal runtime streams selected experts from SSD and exposes an app, CLI, and loopback API.
An open-source Swift and Metal runtime that runs the Gemma 426B A4B model on Apple Silicon Macs, including machines with 8 GB of memory. It streams model experts from SSD and includes a Mac app and CLI.
Runs the Gemma 426B A4B model locally on Apple Silicon Macs by streaming model experts from SSD, with a Mac app and CLI.
2 in the library.