AI product Open source
Speech To Speech is an open-source, modular voice-agent pipeline from Hugging Face. It chains voice activity detection, speech-to-text, a language model, and text-to-speech as separate stages connected by queues; each stage has interchangeable backends. The language-model stage supports OpenAI-compatible protocols and can use hosted providers, Hugging Face Inference Providers, vLLM, or llama.cpp, while the system streams transcripts, generated text, tool calls, and synthesized audio.
The project exposes the core OpenAI Realtime event set through WebSocket and WebRTC and includes commands for running a realtime server, a microphone-and-speaker client, or a local setup. It is distributed as a Python package and requires Python 3.10 or newer. The repository says the pipeline is used as the conversation backend for Reachy Mini robots.
1 use taken from transcripts — each links to the moment in the video.
An open-source modular voice-agent pipeline that chains voice detection, speech-to-text, an LLM, and text-to-speech. Its stages are swappable and it provides an OpenAI real-time-compatible API for hosted or local models.
1 in the library.