Skip to content
buildbay.

GitHub stars

50,616
—Collecting 30d data

Stats updated

39,304 stars on Sep 27 to 50,616 stars on Oct 1, up 11,312.5 measured star snapshots

What is VoiceStudio?

VoiceStudio brings voice creation and audio production into one desktop workspace. It combines reusable voices, dubbing, transcription, and long-form narration with a choice of speech engines, so creators can build a workflow around their own machine and developers can connect it to their existing tools.

VoiceStudio is an open-source desktop app for voice cloning, voice design, video dubbing, dictation, transcription, and audiobooks. Creators can produce speech and edit audio with locally installed models, while developers can connect the same tools to other apps through a local API and MCP. Language support and performance depend on the selected engine and hardware.

Core capabilities

  • Clone a permitted voice sample or describe a new voice, then generate speech with a local model.
  • Create timed video dubs and narrated stories or audiobooks in the same desktop workspace.
  • Connect transcription and speech workflows to other applications through a local API and MCP.

Choose a voice and the engine that runs it

A clean reference recording can provide a voice for speech generation, while voice design starts from a written description. The model catalog offers multiple speech and transcription engines. Features, supported languages, and speed depend on the model and available hardware; remote providers and workers are optional.

  • Voice cloning
  • Voice design
  • Local models

Move from a script or video to finished narration

The dubbing workspace combines transcription, translation, speaker voices, and speech timing. Story and audiobook workflows support longer productions, including chaptered books and multiple voices. A shared desktop app keeps these tasks alongside the voice library and model management.

  • Video dubbing
  • Audiobooks
  • Multiple voices

Use speech tools beyond the studio window

A floating dictation widget brings speech input to desktop work. Developers can use file transcription, streaming text, and native dictation controls through the local speech service. HTTP, WebSocket, and MCP interfaces let editors, scripts, and coding agents reuse these capabilities without opening every workflow manually.

  • Dictation
  • Transcription
  • Local API
  • MCP

Where it fits

Use cases

  1. 01

    Produce a voiceover with a consistent voice

    Start with a recording you have permission to use, or design a voice from a text description. Enter a script and generate speech with the selected engine for a tutorial, presentation, or narrated video.

  2. 02

    Dub a video for another audience

    Work through transcription, translation, speaker selection, and timed speech generation to produce a dubbed version of a video. Available languages and voice features vary by model.

  3. 03

    Turn a long script into an audiobook

    Bring a script or EPUB into the audiobook workflow, organize the narration into chapters, and use the story tools for productions with multiple voices.