Skip to content
buildbay.

GitHub stars

3,119
—Collecting 30d data

Stats updated

3,113 stars on Sep 27 to 3,119 stars on Oct 1, up 6.5 measured star snapshots

What is Willow?

Willow turns compatible ESP32-S3-BOX devices into voice endpoints for a locally administered home. The device handles wake-word detection and audio capture, then routes speech through configured local or self-hosted services before sending commands to Home Assistant, openHAB, or custom HTTP endpoints. The design separates inexpensive room hardware from heavier inference work.

Willow is an open-source, local, self-hosted voice-assistant alternative for the home. Its self-hosted inference server supports language tasks including speech-to-text, text-to-speech, and LLM workloads.

Core capabilities

  • Run wake-word detection and a voice interface on supported ESP32-S3-BOX devices
  • Use self-hosted speech recognition, synthesis, and language inference services
  • Control Home Assistant, openHAB, or custom HTTP endpoints with local commands

Small device and inference server

The ESP32 client provides the room-facing display, microphones, wake-word handling, and interaction loop, while the optional Willow Inference Server runs heavier speech-to-text, text-to-speech, and language workloads on separate local hardware. This split keeps the endpoint compact without requiring cloud inference.

  • ESP32
  • wake word
  • speech
  • self-hosting

Home control and local commands

Official integrations target Home Assistant and openHAB, and local command support can send recognized requests to configured HTTP endpoints. That makes Willow useful both as a smart-home interface and as a general voice front end for private services that expose a small command API.

  • Home Assistant
  • openHAB
  • HTTP
  • local commands

From wake word to visible result

In inference-server mode, a wake word starts recording and the device streams speech until voice-activity detection marks the end. The server returns recognized text, Willow sends it to the configured command endpoint, and the device can present a tone, spoken response, recognized text, and endpoint output. Local-command mode can recognize a configured command on the device and follow the same endpoint-and-feedback path without streaming speech.

  • wake word
  • speech recognition
  • command endpoint
  • feedback

Where it fits

Use cases

  1. 01

    Add room-level voice control to Home Assistant

    Place supported ESP32 voice hardware in a room, configure a Home Assistant endpoint and entities or services, and send spoken home-control requests without making a commercial speaker ecosystem the command broker.

  2. 02

    Keep speech processing on private infrastructure

    Run the companion inference server on local hardware for speech recognition, text-to-speech, and language-model tasks, then direct Willow clients to that service instead of a mandatory public voice API.

  3. 03

    Build a custom spoken command surface

    Map recognized commands to local HTTP endpoints for applications beyond the built-in home integrations, using the voice device as a front end while keeping the action logic in services the operator controls.