AI product Open source

oMLX

oMLX is an open-source LLM inference server for Apple Silicon Macs, managed through a macOS menu-bar application. It serves text, vision, OCR, embedding, and reranking models through compatible API endpoints, with model discovery and management, monitoring, continuous batching, and caching.

View repository Visit site Mentioned in 2 videos ↓

Overview

Its caching system keeps KV cache in a hot in-memory tier and a cold SSD tier, allowing cached context to remain reusable across requests even when the conversation context changes. The project can be installed as a macOS app, through Homebrew, or from source, and includes a CLI for controlling the app-managed server. The repository states that it requires macOS 15 or later, Python 3.11–3.13, and Apple Silicon hardware.

What oMLX is used for

2 uses taken from transcripts — each links to the moment in the video.

  • An open-source LLM inference server for Apple Silicon Macs, managed from the macOS menu bar. It serves text, vision, OCR, embedding, and reranking models through compatible API endpoints and includes caching, batching, monitoring, and model management.

  • A native macOS inference server built on Apple's MLX for serving local models on Apple Silicon. It persists attention-cache blocks to RAM and SSD, supports concurrent requests and multiple model types, and provides OpenAI- and Anthropic-compatible endpoints.

Videos mentioning oMLX

2 in the library.