AI product Open source

VoiceStudio

VoiceStudio is an open-source, fully local desktop application for voice cloning and design, text-to-speech, transcription, dictation, video dubbing, and long-form audio production. It runs on macOS Apple Silicon, Windows, Linux, or Docker and keeps voices, projects, settings, and generated files on the machine by default.

View repository Visit site Mentioned in 1 video ↓

Overview

The application combines a Tauri desktop shell, a React and Vite interface, and a local FastAPI backend with registries for multiple TTS and ASR engines. Its workflows can transcribe and translate video, preserve speaker assignments, synthesize replacement speech, render audiobook chapters, separate vocals, diarize speakers, process batch jobs, and route work to CPU, CUDA, Apple Silicon MPS/MLX, ROCm, or optional remote workers. It also exposes local REST, SSE, WebSocket, OpenAI-compatible audio, and MCP interfaces, including speech synthesis and transcription endpoints.

VoiceStudio is licensed under AGPL-3.0; downloaded models and tokenizers retain their own licenses. The local workflow requires no account, API key, subscription, or usage meter, although remote workers, configured external ASR endpoints, analytics, and other network-backed functions are opt-in. The repository describes the software as an active beta and distributes packaged desktop releases as well as source code.

What VoiceStudio is used for

1 use taken from transcripts — each links to the moment in the video.

  • A fully local, open-source alternative to ElevenLabs for voice cloning, voice design, dubbing, dictation, transcription, and audiobook creation. It runs on a user's machine without accounts, API keys, or cloud services.

Videos mentioning VoiceStudio

1 in the library.