TGSpeechBox v3.10 Beta 9 — the jaw opens, the nasals are nasal, and the lock screen speaks
Beta 8.02 fixed one cliff. Beta 9 is the two-week campaign that came
after it, and it is a real engine change, which is why it gets a number
of its own instead of a point.
Same two rules as before, one of them stricter now: nothing ships on one
person's ears, and nothing ships on ears alone. Every acoustic change
below names its witnesses — code, instruments, blind native listening —
and where one of the three is still owed, the notes say so instead of
pretending. Where a fourth witness existed (period silicon, a second
model's audit) we used it.
The "clenched" voice opens its jaw (#113, #121)
The complaint that started Beta 8.01 — "like a Turkish person learning
Spanish, clenching his teeth" — had a second cause underneath the vowel
one. Our vowels ran their first-formant bandwidths at roughly half of
what real speech, and every reference engine we could measure, uses.
Narrow bandwidth means a resonator that rings and takes time to charge:
the vowel arrives late and holds on tight, and consonants get smeared
into it. That is "clenched."
Every vowel's F1 bandwidth is now 1.8× wider (capped, so the open vowels
don't wash out). Three witnesses, none of which knew about the others:
- Silicon. The Philips PCF8200, the formant chip inside the BraiLab and
the Ciber232 talkers, offers eight fixed F1 bandwidth codes. Frames
decoded from real BraiLab speech chose the codes our new values
quantise to. Rendered on a software model of the chip itself, the
wider bandwidths were the ones that sounded like the genuine article. - Instruments. Every vowel keeps a distinct F1 peak, the F1–F2 valley is
intact, no new seams at any rate from half speed to triple, and
whisper-transcription intelligibility flat to positive with the
English minimal pairs unmoved. - Ears. A blind round on #121 with shuffled, unlabelled pairs: two
native Spanish listeners chose the wider voice on all twelve items
each, 24 for 24. Our Brazilian listener heard no difference on his
set, and that is recorded exactly like that.
What it sounds like: looser, a touch more treble, and — the part that
matters — clearer consonants, because the vowels stop eating them.
Spanish rr gets its contacts back (#115)
Beta 8.01 had a regression found within hours: the trill in "perro"
had been flattened into a tap, because the trill phoneme carries both
flags and a new tap rule grabbed it. Trill now wins. This one shipped
in the repo the same day it was reported; Beta 9 is the first build
that carries it.
Letters get their names on every platform (#122, #123)
Two bugs hiding behind one report about the Spanish letter "r":
- The letter-name dictionary matched exact characters only, so a
capital letter, or an accented one typed on a platform whose C library
only folds ASCII, silently fell through to eSpeak. Case folding now
covers Latin, Latin Extended, Greek and Cyrillic, on every platform,
not only NVDA. - A letter next to punctuation ("r?", "ó.", "¿ñ?") never matched at all.
It does now, and the punctuation stays where it was.
With that in place: Spanish "r" reads as "ere" (native call), and
Brazilian Portuguese gets its own letter names — "a agudo", "a til",
"cê cedilha" and friends.
Brazilian Portuguese: -ão is a diphthong again, and the nasals are nasal (#123)
edu-fblind's report was unusually precise: typing raw "pãw" sounded
right, every real spelling didn't. He was pointing straight at the
cause. The shared Portuguese rules had rewritten eSpeak's nasal-a plus
nasal glide into two full nasal vowels back to back, so "pão", "irmão",
"coração", "não" and "botão" were all two syllables of nasal vowel with
no offglide. Brazilian Portuguese now ends them on a real glide. Blind
round: 11 of 12 in favour of the new rendering.
The twelfth was a fast-rate sentence, and the audit found why: the new
offglide is a semivowel, and a semivowel gets half a vowel's time, so
at 1.6× each -ão kept about 23 ms of glide and "São" less than that. A
new engine knob scales exactly the glide that follows a nasal vowel;
Brazilian gives it 1.6× (about 37 ms at 1.6×, 60 at normal speed), and
the onset "w" is untouched. Five rates, no new seams.
Then the part that was on us. Every Portuguese nasal vowel, and French
"an", had its nasal zero sitting exactly on its nasal pole. In a Klatt
synthesizer that is the "nasalisation off" position: the two cancel,
and what you heard as a nasal vowel was oral vowel shape and slightly
less loudness. Edu said the nasals "stay the same". They were. They now
carry a real nasal pole-zero pair — low nasal resonance up, first
formant down, the direction the measurements of real nasal vowels go.
European Portuguese shares these vowels, so it gets this too; its own
-ão rules are unchanged. French "an" takes the values its siblings
"in" and "on" already had.
And "agora", which read as "agota": a tap with flat voicing and a
little noise is a weak "t". Brazilian taps now get the intensity dip
that Spanish taps have had since Beta 8.01.
Honest note: the glide length and the nasal coupling shipped on code
and instruments, not on native ears, because we wanted this beta to
carry the audit's work rather than sit on it for another cycle. The
blind confirmation zip is on #123. If Edu's ears say we went the wrong
way on any of it, a point release reverts it.
Taps keep one shape across rate, and English flaps get a dip
Found by an outside audit (GPT-6 "Astra", reading the code and
rendering frames): a tap or flap — Spanish "pero", Portuguese "agora",
American English "ladder", "accessibility" — took a different rendering
path in a narrow window of speech rates, roughly 1.5× to 1.75×, than at
every other rate. In Spanish the tap's intensity dip was applied twice
there. The window is closed; the shape natives approved at normal rate
is now the shape at every rate. Verified frame-for-frame at 25 rates:
identical everywhere outside the window, continuous inside it. A blind
English listen at the affected rates heard no difference, which is the
right answer for a continuity fix.
Separately, American English flaps now get an intensity dip of their
own, milder than the Spanish one. Without it a flap is a faint vowel
with a little noise laid on top of it. Blind triplets, English ears:
more of a thump, better at fast rates, and kept at slow ones.
Hungarian: gy and ty between vowels are palatal now
"Fagyi" sounded like an American saying it, with a hard g. The frames
agreed: between vowels, gy was a 26 ms murmur and a 6 ms burst, 32 ms
of consonant, and its closure sat at F2 2200 / F3 2850, which is the
velar pinch that makes a g. Real Hungarian palatals keep F2 and F3
apart. The change: a 40 ms closure, a 24 ms body with a small burst and
a half-voiced palatal release, and the closure formants opened to
2350 / 3200. Word-initial, word-final, geminate and cluster gy and ty
are untouched. Witnesses: the real BraiLab's own "fagyi" for the timing
of the event (its gy is a long, deep loudness dip, not a click),
Eloquence-lineage and qlatt duration tables for how long a voiced stop
between vowels should be (80 to 105 ms; ours was 32), whisper on a
sixteen-word battery (four gy words newly spelled right, none worse),
and a blind six-pair round on a native ear, five for the new gy, with
the fast rates checked separately. The single value that turned "starts
on a hard g" into "can't complain" was the formant pair, not the burst.
Found on the way: a Hungarian rule meant for geminate stops used a flag
name the engine does not know, so it matched every repeated sound. Inside
a word nothing showed; across a word boundary it silenced the first of
two identical fricatives: "kis sas" and "lesz szép" lost a consonant.
Fixed.
A new pitch mode: Arató (BraiLab)
The BraiLab talking computers of the late 1980s, designed by Arató
András and Vaspöri Teréz, gave every punctuation mark its own melody so
a blind reader could hear from the intonation alone how a line ended,
even when the next line was already on its way. Arató described the
design in his 1992 dissertation. Tonight the original BraiLab PC
program, running in an emulator, gave us its melodies as numbers: the
start pitch of every clause and the pitch step of every 12.8 ms frame.
Then we went further: the 1991 program was carefully studied and
referenced, and it turned out the melodies are not stored anywhere. One
routine writes them, clause by clause, from eighteen constants, and the
program written nine years later still carries the same routine
unchanged. The new mode is a port of that routine. Statements start low after "a" or "az" and
lift on the next word, decline in a straight line and plunge at the end.
Yes/no questions creep up and rise sharply on the syllable before the
last, then fall; a two-syllable question has its own shape. Wh-questions
and exclamations start high and fall from a fixed point. A comma clause
falls to its last word and rises over it, then the next clause starts
fresh. The question words are recognised the way the original did it, by
the first letters of the sentence. Checked against the emulated program
on sixteen sentences covering every branch, the contours match within a
few hertz at normal and double speed, and one native ear compared
same-sentence pairs and heard the same melodies.
It is selectable on every platform as "Arató (BraiLab)"; no language
switches to it by default in this beta. Every constant is a pack
setting, so other languages can be measured and tuned the same way.
Android: speaks on the lock screen after a reboot
Direct Boot support, ported from Panthera-speech. Android hides speech
engines that can't run before the phone's first unlock, so after a
reboot TalkBack fell back to Google's engine on the lock screen. The
engine's data and settings now live in device-protected storage and
the service declares it can run there. Your saved voice and language
are carried over once; the first launch after this update re-extracts
the language data (about 26 MB) into the new place.
Also in this release
- SAPI 5 defaults to 22050 Hz on a fresh install, matching every other
platform (#117, gregodejesus2). - The DSP documentation now tells the truth about tilt: on the voice
source, negative is brighter; on aspiration and frication, negative is
darker; and the "Voice tilt (brightness)" slider going UP means
darker on purpose — think of a valve you are closing. - Test suite: a tripwire pins the Python frame mirror to the C++ header
so the two can never drift again (the Beta 1 bug, structurally
prevented); the Beta 8 clarity regression pins are committed; test
renders land in their own folder instead of the repo root. - Installer README describes the installer that actually ships.
What we found and did not ship yet
English words inside Russian or Bulgarian text (#118): the language of
each span is dropped in one place on every platform, and the fix is one
mechanism for all of them. It is structural, and it is the next thing
on the bench.
The "dense" voice, finally measured. We compared the neutral voice
against an ETI-Eloquence-lineage reference render, pitch-matched and
duration-matched, expecting to find it bottom-heavy. It isn't: the
long-term spectrum agrees within a decibel. What differs is the
loudness contour. The reference's typical speech frame sits 9 dB under
its loudest, ours 4; its syllable-rate loudness swing is nearly twice
ours. Our frontend writes almost every voiced sound at the same level
and the engine clamps voice amplitude at its ceiling, so a syllable
has nowhere to rise to. Fixing that is an amplitude model plus headroom
in the engine, with instruments and blind rounds around it, and it is
the headline of the next beta.
Credits: gregodejesus2 and 29-Bloo for the blind Spanish rounds and the
letter reports; edu-fblind for the most precise Portuguese bug report
this project has received; and GPT-6 Astra for an audit whose every
claim survived verification, and for catching two mistakes of ours in
the process.
— Tamas + Claudeo (Fable 5.1)
Linux builds (x86_64 + aarch64) auto-generated from tag v-310b9.
Extract and run ./install.sh to install, or use the tgsp wrapper directly.