github tgeczy/TGSpeechBox v-310b901
TGSpeechBox 3.10 beta 9.01 — hotfix

pre-release5 hours ago

TGSpeechBox 3.10 beta 9.01 — hotfix

This is an urgent hotfix to beta 9, based on what people found in the first hours after the release, plus one change to the English voicing to make it clearer. If you are on beta 9, take this one.

The fixes. The breath that hung in the air after a long pause, most noticeable when a screen reader lands on something like "DoubleTalk" and then goes quiet, is gone. The engine deliberately keeps a little of a released stop's aspiration alive so the release has a tail, but it kept it for the whole length of whatever silence followed, so a two-second pause was a two-second hiss at low level. The tail now decays over the first 150 ms of any silence and the rest is true silence, on every platform. And the Arató pitch mode could split a gliding vowel in two and play its offglide twice ("why-i", "voi-oice"); it now leaves diphthongs whole. Both from issue #125, thank you for the fast report.

The voicing change, US English only. Until now every voiced sound was one loudness held flat for its whole length, joined to the next by a very short crossfade, and nearly every voiced sound in the packs had the same source amplitude, so a nasal was as loud as a vowel and "the" was as loud as the word after it. That is a large part of what made the voice sound spliced and gritty behind the words. A new amplitude contour gives each sound its own level and shape the way speech has it: stressed vowels start strong and settle a little, unstressed ones sit lower, nasal murmur decays toward the closure, voiced fricatives dip into a real valley, voiceless consonants sit a touch down, every boundary is a 20 ms slope instead of a step, and each clause slopes down gently toward its last word. It was tuned over five listening rounds and a blind pair test against ETI-Eloquence-lineage reference renders, at 1x, 2x and 3x, and the pack's output gain is raised so the overall loudness matches beta 9 within a decibel. Under the hood this is a DSP version bump (v9): two new per-frame parameters, a voice-amplitude end target and an onset glide time, with the same forward and backward compatibility as every earlier extension. The other English variants, Hungarian and every other language render exactly as in beta 9, in both pitch modes, and keep whatever settings you had; the same contour can be enabled per language from the pack files once it has been listened to there. The amplitudeContour settings are documented in Tuning.md and the new frame fields in Developers.md.

Windows, NVDA add-on, SAPI, phoneme editor, Android and Linux builds are attached; the Play Store and TestFlight builds follow. As always, if something sounds off on your system, say so in the issues; the contour is a pack setting and can be turned down or off per language.

Built with Claude (Anthropic) as engineering partner.


Linux builds (x86_64 + aarch64) auto-generated from tag v-310b901.

Extract and run ./install.sh to install, or use the tgsp wrapper directly.

Don't miss a new TGSpeechBox release

NewReleases is sending notifications on new releases.