Speaches
speaches-ai · API tools
GitHub stars
3,693Stats updated
What is Speaches?
Speaches packages open speech models behind endpoints shaped for existing OpenAI clients. A self-hosted deployment can transcribe or translate recordings, stream partial transcripts, synthesize speech, run real-time voice sessions, and load or unload models dynamically on CPU or GPU infrastructure.
Speaches is an OpenAI API-compatible server for streaming transcription, translation, and speech generation. It supports CPU and GPU deployments, and can be deployed with Docker or Docker Compose.
Core capabilities
- Transcribe and translate audio with streaming OpenAI-compatible endpoints
- Generate speech with Kokoro and Piper models
- Run real-time voice, diarization, model management, and voice activity detection
Speech services behind familiar endpoints
Speaches implements OpenAI-shaped audio endpoints for transcription, translation, and speech generation, plus APIs for model discovery and management. Existing SDKs can target the server while faster-whisper, Kokoro, and Piper handle the underlying speech workloads.
- OpenAI compatible
- speech-to-text
- text-to-speech
Streaming and real-time audio workflows
Transcriptions can stream before an uploaded recording finishes processing, and the Realtime API supports conversational and transcription-only WebSocket sessions. Voice activity detection, diarization, dynamic model loading, CPU and GPU images, and Docker deployment round out the self-hosted operating surface.
- streaming
- Realtime API
- diarization
- Docker
Where it fits
Use cases
- 01
Replace a hosted speech endpoint
Point compatible clients at a self-hosted Speaches base URL to transcribe, translate, or synthesize audio while keeping model execution on infrastructure selected by the operator.
- 02
Stream live transcription
Receive incremental transcription results over server-sent events or use the Realtime API in transcription-only mode for live subtitles, meeting notes, voice capture, and accessibility workflows.
- 03
Prototype a real-time voice assistant
Combine speech recognition, an external conversation model, and speech synthesis over the OpenAI-compatible WebSocket interface to test interactive voice applications with configurable local audio models.