AI product Open source
Backburner is a pre-release local-LLM inference tool by StayLameBro that uses an iPhone alongside an Apple Silicon Mac to run Qwen3.8-27B over a USB-C connection. Its llama.cpp-based engine splits prefill computation between the Mac and phone: the Mac runs the initial model layers while the iPhone runs later layers, and for contexts beyond the Mac's available 8-bit capacity the phone stores older key-value pages and computes attention over them. The phone can also use its GPU and Neural Engine during decoding, while the Mac uses custom Metal, SME2, and DFlash2 speculative-decoding kernels. It exposes an OpenAI-compatible server at localhost:8080 and supports iPhones 15 Pro or newer and M-series iPads, although the documented testing uses iPhone 16 Pro Max and 17 Pro Max devices. The project requires a 10 Gb/s USB-C cable, distributes its iPhone app through AltStore or an Xcode build, and is released under the MIT license.
1 use taken from transcripts — each links to the moment in the video.
Uses an iPhone alongside a Mac running a local model to process long prompts or hold older context that does not fit comfortably on the Mac. The devices communicate over a USB-C cable.
1 in the library.