AI product Open source
Splash is a local inference engine for Apple silicon developed by Inco AI. It serves selected language models to coding agents and clients compatible with the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages APIs, with streaming, tool calls, JSON Schema output, image input, and inline PDF support.
Its runtime uses model-specific fused Metal kernels, weights packed for those kernels, a dedicated DFlash 2 draft model for speculative decoding, and a startup-computed memory plan for context, KV cache, and batching. The served model packages include the target model and its draft model; plain MLX or Transformers checkpoints are not supported. Splash runs locally on macOS, binds to a local HTTP server by default, and can be installed through Homebrew. The project is licensed under Apache-2.0, while model weights retain their own licenses.
2 uses taken from transcripts — each links to the moment in the video.
Runs language models locally on Apple silicon using model-specific kernels and draft models for speculative decoding, exposing OpenAI- and Anthropic-compatible APIs.
A local inference engine for Apple silicon that serves OpenAI- and Anthropic-compatible APIs, supports vision and tool calls, and connects to coding agents.
2 in the library.