A pre-release for field validation. HACS offers this only to users who have enabled beta versions for this repository. It is not a public release, and the changes below still ship to everyone in the next monthly bundle as 3.43.0.
It's cut so @andymcmanus can check #671 on a fully local Voice PE + Kokoro + qwen3 setup: the satellite is silent for the whole of a tool-using turn, then says everything at once. It supersedes v3.43.0-beta.4 and still includes beta.4's Ring event-recording fixes for #491.
What changed since beta.4
HGA's text-to-speech engine now streams (#679). The Assist pipeline only passes a reply to TTS while it is still being written if the engine supports streaming input, and HGA's didn't. Now:
- The reply is spoken a sentence at a time. The first sentence is synthesized on its own; after that, one request covers every sentence that arrived in the meantime. Pieces are joined as mp3, and Home Assistant converts them for the satellite.
- A finished sentence followed by a 0.4 s pause (the agent stopping to call a tool) is spoken during the pause, not held until the answer.
- A message that arrives whole (
tts.speak, announcements) is still one request in the player's preferred format. - Reasoning that a model writes inline as
<think>…</think>is removed from the streamed text so it is never spoken. Ollama models configured through HGA were not affected.
New option: Spoken acknowledgement before tool use (#680). In Global Options (Configure) → Spoken acknowledgement before tool use. When the model calls tools without saying anything first, the phrase is spoken the moment the tools start. It applies only on Assist satellite turns, at most once per reply. Empty (the default) means off. A phrase without a closing sentence mark gets a full stop.
What to validate
Pipeline: Text-to-speech = HGA's TTS provider. Streaming depends on it; another engine that does not stream will still speak the reply all at once.
- Plain reply, no tools. A longer answer should start speaking after its first sentence rather than after the whole reply.
- Tool-using "why" question, option empty. If the model writes anything before its tool call, that should be spoken while the tools run. qwen3 usually writes nothing first, so expect silence until the answer starts. It should still start sooner than on beta.4.
- Same question with the option set to
Let me check.The phrase should be heard within about a second, then the answer. - Anything odd: words run together, a clipped or repeated sentence, a gap or click between sentences, reasoning text read aloud. Also check that
tts.speakannouncements still sound right.
Timings as before (question end → first audio → answer audio) would be ideal. Logs for custom_components.home_generative_agent.tts: debug help if something goes wrong.
Everything else already on main since v3.42.0 is included. See CHANGELOG.md under Unreleased.