AI product Open source

Speech To Speech

Speech To Speech is an open-source, modular voice-agent pipeline from Hugging Face. It chains voice activity detection, speech-to-text, a language model, and text-to-speech as separate stages connected by queues; each stage has interchangeable backends. The language-model stage supports OpenAI-compatible protocols and can use hosted providers, Hugging Face Inference Providers, vLLM, or llama.cpp, while the system streams transcripts, generated text, tool calls, and synthesized audio.

View repository Mentioned in 1 video ↓

Overview

The project exposes the core OpenAI Realtime event set through WebSocket and WebRTC and includes commands for running a realtime server, a microphone-and-speaker client, or a local setup. It is distributed as a Python package and requires Python 3.10 or newer. The repository says the pipeline is used as the conversation backend for Reachy Mini robots.

What Speech To Speech is used for

1 use taken from transcripts — each links to the moment in the video.

  • An open-source modular voice-agent pipeline that chains voice detection, speech-to-text, an LLM, and text-to-speech. Its stages are swappable and it provides an OpenAI real-time-compatible API for hosted or local models.

Videos mentioning Speech To Speech

1 in the library.