VoiceStudio
debpalash · Audio
GitHub stars
50,616Stats updated
What is VoiceStudio?
VoiceStudio brings voice creation and audio production into one desktop workspace. It combines reusable voices, dubbing, transcription, and long-form narration with a choice of speech engines, so creators can build a workflow around their own machine and developers can connect it to their existing tools.
VoiceStudio is an open-source desktop app for voice cloning, voice design, video dubbing, dictation, transcription, and audiobooks. Creators can produce speech and edit audio with locally installed models, while developers can connect the same tools to other apps through a local API and MCP. Language support and performance depend on the selected engine and hardware.
Core capabilities
- Clone a permitted voice sample or describe a new voice, then generate speech with a local model.
- Create timed video dubs and narrated stories or audiobooks in the same desktop workspace.
- Connect transcription and speech workflows to other applications through a local API and MCP.
Choose a voice and the engine that runs it
A clean reference recording can provide a voice for speech generation, while voice design starts from a written description. The model catalog offers multiple speech and transcription engines. Features, supported languages, and speed depend on the model and available hardware; remote providers and workers are optional.
- Voice cloning
- Voice design
- Local models
Move from a script or video to finished narration
The dubbing workspace combines transcription, translation, speaker voices, and speech timing. Story and audiobook workflows support longer productions, including chaptered books and multiple voices. A shared desktop app keeps these tasks alongside the voice library and model management.
- Video dubbing
- Audiobooks
- Multiple voices
Use speech tools beyond the studio window
A floating dictation widget brings speech input to desktop work. Developers can use file transcription, streaming text, and native dictation controls through the local speech service. HTTP, WebSocket, and MCP interfaces let editors, scripts, and coding agents reuse these capabilities without opening every workflow manually.
- Dictation
- Transcription
- Local API
- MCP
Where it fits
Use cases
- 01
Produce a voiceover with a consistent voice
Start with a recording you have permission to use, or design a voice from a text description. Enter a script and generate speech with the selected engine for a tutorial, presentation, or narrated video.
- 02
Dub a video for another audience
Work through transcription, translation, speaker selection, and timed speech generation to produce a dubbed version of a video. Available languages and voice features vary by model.
- 03
Turn a long script into an audiobook
Bring a script or EPUB into the audiobook workflow, organize the narration into chapters, and use the story tools for productions with multiple voices.