AI product Open source · MIT

Voicebox

Voicebox is a local-first, open-source AI voice studio developed in the jamiepine/voicebox repository. It combines voice cloning, text-to-speech, speech-to-text dictation, voice profiles, audio effects, and agent voice output in a desktop application.

View repository Visit site

Overview

It generates speech through seven switchable TTS engines, including Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox, HumeAI TADA, and Kokoro, and supports voice cloning from reference audio, preset voices, multilingual generation, automatic text chunking, crossfading, and multi-track story editing. Whisper-based transcription powers global dictation, in-app recording, and the transcription API; an optional local Qwen3 LLM can refine dictation or rewrite text according to a voice profile's persona.

The application exposes REST endpoints for speech generation, agent speech, transcription, and profile management, plus a built-in MCP server with tools for speaking, transcribing, and browsing captures and profiles. It runs models and stores voice data locally, using Tauri with a React/TypeScript frontend and FastAPI backend; platform backends include MLX on Apple Silicon and PyTorch on CUDA, ROCm, XPU, DirectML, or CPU. The repository is licensed under the MIT License and provides macOS, Windows, and Docker distribution, with Linux build-from-source instructions.