TGSpeechBox v3.10 Beta 8.01 — the vowels land, and the natives judged
Beta 8 shipped our biggest Spanish push ever, and within hours our
sharpest-eared tester told us something was off: "like a Turkish person
learning Spanish, clenching his teeth." He was right, and finding out
WHY he was right turned into the most instructive week this project has
had. This release is that week.
Two rules governed every change below: nothing ships on one person's
ears again, and nothing ships without a measurement. Every fix here was
verified three ways — acoustic analysis of rendered audio, an automatic
transcription referee, and blind A/B listening tests scored by native
Spanish speakers who did not know which clip was which. Across two
rounds and two listeners, the fixes won the blind vote 39 times out of
39.
Vowels land on their targets (#113)
The root cause of "clenched teeth" was humbling: a coarticulated vowel
in our engine spent its ENTIRE duration in transition — ramping from
its onset toward its exit — and never actually rendered its own target.
The /o/ in "dos" swept from 1129 to 1105 Hz without ever touching its
910 Hz home: acoustically, that's not /o/ with transitions, it's a
German ö. Beta 8's coarticulation work (correct in direction, and the
consonant half of it measurably improved intelligibility) made this
old structural gap audible everywhere.
Vowels now render the way reference formant synthesis and natural
speech both do: a short transition in, a genuinely held steady state at
the vowel's canonical formants, and a transition out. Spanish opts in
this release; other languages are byte-identical, and that's verified,
not assumed. We did try the same flag on English tonight before
shipping: the transcription referee scored it a wash and a careful
English ear heard no difference — which is itself a finding. English's
"blendy" character is NOT the ramp disease; it lives somewhere else,
and it gets the full Spanish-style decomposition next cycle. Hungarian
joins that queue with a diagnosis already tested and refined tonight:
its palatal stops are position-dependent — word-final /ɟ/ ("nagy") has
a real closure and release, while word-initial and intervocalic
("gyerek", "fagyi") render a near-zero closure and a fricated puff
instead of a burst. That's the same "events, not blurs" surgery the
Spanish soft consonants got in beta 8, aimed at Hungarian's palatals,
with a period-correct Hungarian reference synth on the analysis bench.
The little vowel around the R (#113's second gem)
Our tester's second clue — "for, far, fer, fir, fur" — looked like
example words. It was a diagnosis. All five rendered with the SAME
vowel, and had for months, in every build. Spanish inserts a tiny
vocalic element around taps and trills (the "svarabhakti" vowel — real,
documented phonetics), and ours was one fixed neutral schwa no matter
which vowel it belonged to. Every r-final word in Spanish ended in the
same blur: "far" was literally rendering as "far-uh".
The fix is a new frontend pass — the svarabhakti pass — and it's
substantial work we're a little proud of: a per-language rule stage,
built directly on the peer-reviewed phonetics rather than on vibes.
What the literature says, it now does:
- The inserted vocoid replicates the quality of the syllable's nuclear
vowel — same articulatory gesture, not an independent schwa — at
roughly a quarter of the nuclear vowel's duration (da Silva 2024,
Journal of Speech Sciences). - The inheritance runs in the carryover direction — the preceding
vowel colors through the tap — per Recasens' electropalatographic
studies of tap/trill coarticulation (Recasens 1991; Recasens &
Pallarès 1999, Journal of Phonetics). - The tap itself takes on its neighbor's formant color, and the tap's
amplitude dip — the intensity difference against the flanking vowels
that IS the primary acoustic cue of the Spanish tap (Perry et al.
2023 & 2024, JASA) — is restored to carry the rhotic percept.
The pass is fully tunable per language pack (svarabhaktiInheritEnabled,
vowel weight, duration/amplitude scaling, tap blend and dip — documented
in Tuning.md), ships enabled for Spanish, and is byte-identical for
every pack that doesn't opt in. "amor" ends in /o/. "primero" has its
R. Five words are five words.
Also fixed in the same dig: the single-word final hold used to stretch
a word-final tap (a 15 ms tongue flick cannot be held — holding it
manufactured a schwa syllable); the hold now lands on the vowel where
citation speech puts it.
Australian English: "thousand" is one word again
The AU MOUTH diphthong glided from an open onset to an offglide that
was barely closer — F1 moved 34 Hz — so it read as two separate open
vowels: "th-ou-au-sand", drawn out and washy. The offglide now closes
properly (while keeping the broad-Australian rounding), and the
diphthong is a single gesture. Doctest-pinned like everything else.
NVDA: volume installs at 100% (#114)
The NVDA driver mapped its volume slider on a different scale than
every other platform — its 100% was internally "a third louder than
everyone else's", so fresh installs parked partway down the slider.
Fresh installs now report 100% like Android, Linux and SAPI, and if you
already saved a volume, it is rescaled once automatically: nothing
changes audibly, only the number.
Housekeeping with teeth
Android: beta 8's APK shipped with a stale data-version marker, which
quietly kept beta 7's language packs running under beta 8's engine on
updated devices. This build carries the corrected marker: your device
re-extracts once and engine and data agree again. (Release-checklist
rule added so it cannot happen twice.)
iOS: a real fix for pause mode (it was structurally inert — VoiceOver's
requested pauses were being collapsed to near-zero) is already in the
tree and arrives with the next TestFlight build once the Mac has done
its part.
The voice, round two: the head un-squishes
Beta 8 gave the voice its presence stage. Tonight, working live with
ears on headphones, the July "voice identity" elimination board finally
got its missing entries — and three keepers shipped:
- F4/F5 placement, +6% across the whole inventory. The vocal-tract
axis that separates chest-dominant from head voice in the classic
synths. +6% was the keeper; +12% overshot into "head squished too
high", which told us we were on the real lever. - The presence stage is finished: a second +4 dB peak at 4.9 kHz
closes the last measured gap against the reference-lineage spectra. - The speaker sits closer: two glottal-source constants trimmed
(open-phase F1 damping, source-tract darkening) — reads as proximity,
not EQ.
Tested and rejected, for the record: default-pitch changes (transposes
the voice, cures nothing) and removing the subglottal "chest" stage
(thins soft Spanish vowels — it stays, it's load-bearing).
Hungarian: gy and ty grow teeth
Hungarian's palatal stops are affricated — and ours released as a 6 ms
puff, so "fagyi", "gyerek", "kutya" went soft and papery while
word-final "nagy" was fine. The palatals now carry a real burst in
initial and intervocalic position, and word-final position keeps the
softer release it already did well (boosting it read as bumpy — a
native ear caught that immediately). fagyi says fagyi. This is the
first bite of a full Hungarian campaign; two period-correct Hungarian
reference synths are already on the analysis bench for the next one.
What we need from you
- Spanish listeners: connected sentences, and r-words at every
position — amor, pero, primero, corchete, arreglado. Two of you
already voted blind; now tell us how it feels in real use. - Australians and AU-voice users: thousand, about, down, house, now.
- Everyone: the voice update is global and deliberate — every language
sounds a touch more head-voiced, closer, and clearer up top. The
PHONETIC changes stay confined to Spanish, Hungarian palatals and the
en-AU MOUTH vowel (that confinement is verified, not vibes). If a
language's phonetics — not its overall timbre — changed on you,
that's a bug report we want.
Credit where due
This release belongs to gregodejesus2 and 29-Bloo (Dani), who between
them turned a complaint into a methodology: blind A/B zips, returned
with verdicts, 39 picks out of 39 for the fixes — and whose minimal
pairs found a bug our instruments had walked past for months. The
volume catch (#114) was Greg's too, tested across all four platforms
before reporting, which made it a five-minute diagnosis. Reference
measurements against renders from the classic commercial formant
synthesizer lineage; peer-reviewed phonetics as the tiebreaker;
faster-whisper as the referee.
— Tamas + Claudeo (Fable 5)