An open-source LLM inference server for Apple Silicon Macs, managed from the macOS menu bar. It serves text, vision, OCR, embedding, and reranking models through compatible API endpoints and includes caching, batching, monitoring, and model management.