Skip to content
buildbay.

GitHub stars

3,693
—Collecting 30d data

Stats updated

3,682 stars on Sep 27 to 3,693 stars on Oct 1, up 11.5 measured star snapshots

What is Speaches?

Speaches packages open speech models behind endpoints shaped for existing OpenAI clients. A self-hosted deployment can transcribe or translate recordings, stream partial transcripts, synthesize speech, run real-time voice sessions, and load or unload models dynamically on CPU or GPU infrastructure.

Speaches is an OpenAI API-compatible server for streaming transcription, translation, and speech generation. It supports CPU and GPU deployments, and can be deployed with Docker or Docker Compose.

Core capabilities

  • Transcribe and translate audio with streaming OpenAI-compatible endpoints
  • Generate speech with Kokoro and Piper models
  • Run real-time voice, diarization, model management, and voice activity detection

Speech services behind familiar endpoints

Speaches implements OpenAI-shaped audio endpoints for transcription, translation, and speech generation, plus APIs for model discovery and management. Existing SDKs can target the server while faster-whisper, Kokoro, and Piper handle the underlying speech workloads.

  • OpenAI compatible
  • speech-to-text
  • text-to-speech

Streaming and real-time audio workflows

Transcriptions can stream before an uploaded recording finishes processing, and the Realtime API supports conversational and transcription-only WebSocket sessions. Voice activity detection, diarization, dynamic model loading, CPU and GPU images, and Docker deployment round out the self-hosted operating surface.

  • streaming
  • Realtime API
  • diarization
  • Docker

Where it fits

Use cases

  1. 01

    Replace a hosted speech endpoint

    Point compatible clients at a self-hosted Speaches base URL to transcribe, translate, or synthesize audio while keeping model execution on infrastructure selected by the operator.

  2. 02

    Stream live transcription

    Receive incremental transcription results over server-sent events or use the Realtime API in transcription-only mode for live subtitles, meeting notes, voice capture, and accessibility workflows.

  3. 03

    Prototype a real-time voice assistant

    Combine speech recognition, an external conversation model, and speech synthesis over the OpenAI-compatible WebSocket interface to test interactive voice applications with configurable local audio models.