TGSpeechBox 3.10 release candidate — clearer consonants, profiles that travel, and the engine in Python
This is the release candidate for TGSpeechBox 3.10, after ten betas since 3.01 (April to September). If nothing serious turns up, it becomes 3.10 in October, unchanged. Most of it came from you: reports in the issues, blind listening rounds, recordings and logs. The theme turned out to be clarity: consonants that land, vowels that sit where they should, and a voice that behaves the same on every platform.
Why now
Honestly, 3.10 stretched on long enough. Some things we hoped for didn't make it in: a currency dictionary (#83), more Croatian testing and fixes, and the others listed under Next. They're well scoped for the next release, and it mattered more to put out a stable release for everyone. With iOS 27 out and a new Android release here too, we went ahead now rather than later.
The voice
- Consonants are clearer everywhere. /s/ no longer smears toward /z/ after a vowel, voiced fricatives like /z/ and /v/ have their buzz, and fricatives carry a little more presence.
- Vowels are looser and less "clenched" (wider first-formant bandwidths), checked in a blind round with native listeners.
- A presence lift on the output of every voice and platform, with slightly raised upper formants, takes away the muffled, sore-throat quality some of you described, and the speaker sounds a little closer.
- Slow rates no longer produce a phantom echo; formants glide instead of stepping.
- Word-final stops are audible ("let", "cat"), and the breath after a released stop fades out in a pause instead of hissing through it.
- Rolled r's start with a contact and end on the release, which tidies the trill in every language that uses one.
- A light breath floor on vowels takes a little of the buzz off without sounding breathy.
- Pauses land in the same places on every platform, and now also at dashes, around parentheses, and before Spanish ¿ and ¡ (#133). Colons and semicolons get a full pause everywhere, and web addresses like example.com no longer split.
- US English has a loudness contour: sounds rise and fall in level the way speech does, so it sounds less spliced.
Languages
- Spanish: /g/ and /d/ between vowels are real events rather than blurs ("digo" no longer sounds like "y hoy"); the vowel around a tap or trill copies its neighbour ("amor" ends in /o/); doubled consonants within a word collapse; "las semillas" no longer restarts its /s/; /sj/ and /ts/ no longer drift toward "sh", and "siguiente" is cleaner; Mexican and Castilian /s/ are brighter; the letter y is "i griega", and r is "ere".
- Brazilian and European Portuguese: -ão ends on a real glide (pão, não, coração), nasal vowels are actually nasal, "agora" no longer sounds like "agota", and the letter names are right, accents included (ê and ô as "ê circunflexo", õ as "ô til").
- Hungarian: gy and ty are real sounds between vowels and at the start of words ("fagyi"), doubled consonants across words are both spoken ("kis sas"), and "arra" is no longer "ah-ra".
- English (UK): word-final stop bursts are no longer too quiet. English (Australia): the vowels in "price" and "mouth" are Australian. US English: flaps have a little more thump.
- Swedish: words ending in rt, rd and rn ("bort", "bord", "barn") no longer lose their last consonant.
- Italian has stronger stress, Dutch a uvular r, French a real nasal "an", Croatian released final stops, open-mid e and a "uvod" that is no longer silent, and Polish ś ź ć dź sit further from sz ż cz dż.
- Portuguese, Swedish and Polish changes were checked with instruments and the phonetics literature more than by native ears; if your language sounds off, tell us which word.
NVDA add-on
- Follows the page's language when NVDA's automatic language switching is on. Text in your own language keeps your own dialect (a Brazilian user on a page marked "pt" stays Brazilian, a UK user stays UK), and switching back to a language already spoken costs about what eSpeak's own voice change does.
- Old settings saved from the NV Speech Player days can no longer overwrite the language files: a leftover "stop closure mode: none" was removing the pause before t, d and b for some of you. If consonants still run together for you, set Stop closure mode back to "Vowel and cluster" once.
- NVDA used to drop Spanish ¿ and ¡ before speech reached any voice. The add-on now brings a small symbol dictionary (NVDA 2024.4 and later) that passes them through, so TGSpeechBox can pause where a question or exclamation opens. It applies to every voice while the add-on is installed; if you'd rather not have it, set "Send actual symbol to synthesizer" to "never" for ¿ and ¡ in NVDA's Punctuation/symbol pronunciation dialog.
- Word-final sounds are no longer cut short, the capital-letter beep lands on the letter, volume 100% means the same as everywhere else, a word read twice sounds the same twice, when you move on quickly, nothing of the item you left leaks into the next one, and words either side of a dash, ¿ or ¡ no longer run together into one.
SAPI
- Long text starts speaking right away (the engine streams while it renders), stopping speech takes effect at once, and an interrupted item no longer delays the next one.
- Words either side of a dash, ¿ or ¡ no longer run together into one.
- Built against eSpeak NG 1.53 from its latest source; defaults to 22050 Hz on a fresh install.
Android
- Speaks on the lock screen right after a restart (Direct Boot).
- Runs on Wear OS watches; Android 16 is the target.
- Fast navigation no longer falls silent, and each utterance starts quicker.
iOS and macOS
3.10 is on its way to the App Store. VoiceOver's speech volume rotor now changes TGSpeechBox's volume (#119; it had no effect before). VoiceOver's pauses are honoured (they were collapsed to almost nothing), stopping mid-sentence is reliable, there is a pause between VoiceOver's separate announcements on iOS 27 ("dock", "page 1 of 4"), scaled by your pause setting, and TGSpeechBox no longer latches onto the wrong eSpeak voice when languages switch quickly.
Linux
Adam's pitch is fixed, speech-dispatcher's default pitch is read correctly, and profiles apply their voice settings. x86_64 and aarch64 builds are attached as before.
Voice profiles and the phoneme editor
- A voice profile sounds the same on every platform. Your own sliders adjust it from there, and a slider back at its default gives you the profile as it was saved. Beth and Bobby use their own written settings everywhere.
- Profiles can set their own inflection, and now their own chorus (chorusDepth and chorusDetuneHz under voicingTone).
- In the phoneme editor, Save to Profile now asks for a name and keeps the pitch and formant shape of the voice you started from, and loading a profile no longer shows it empty (saving one could drop parts of Beth). Phonemes gain frication tilt and closure gap fields. The editor's neutral slider positions now match NVDA's.
- New pitch mode: Arató (BraiLab), a port of the intonation routine of the 1991 Hungarian BraiLab screen reader. Its melody keeps its shape at any pitch.
Letters
- Letter names are read on Android, iOS and Linux too, a lone letter next to punctuation is spelled as a letter, and capitals find their names.
New: the engine in Python
People asked about the DSP behind the voice, so here it is: a small Python package, tgspeechbox, with the same engine every platform runs. You give it IPA, and you get audio back, plus every frame and the frontend's trace of what each processing step did.
pip install tgspeechbox-3.10.0-py3-none-win_amd64.whl (the file for your platform)
python -m tgspeechbox --ipa "həˈloʊ wˈɜːld" --out hello.wav
Wheels for Windows (64- and 32-bit), macOS (Intel and Apple Silicon) and Linux (x86_64 and aarch64) are attached. The package is MIT licensed and contains no eSpeak NG (which is GPL), so it takes IPA rather than text. It will go on PyPI once it has proven itself.
Tests
3.10 is the first release with a test suite behind it: the engine, the language packs, and now the NVDA add-on and the SAPI voice themselves, driven the way NVDA and Windows drive them. Every fix since beta 10 came with a test that failed first.
Since beta 10
For beta testers, what changed since beta 10:
- Portuguese: ê, ô and õ are read with their accents (#130).
- iOS: VoiceOver's speech volume rotor works (#119).
- Pauses at dashes, parentheses, ¿ and ¡, from one clause splitter every platform now shares, and NVDA now passes ¿ and ¡ through (#133).
- Fixed: leftover audio on fast navigation, on SAPI and NVDA (#135); "agota" sounding like "agora" on NVDA (#127); dialect switching (#131, #137); the Arató melody scales with pitch (#136); words either side of a dash, ¿ or ¡ no longer run together; every platform starts each request fresh; and switching back to a language is fast.
- New: chorus in voice profiles (#124), and the Python wheel.
Next
These moved to the next cycle: a currency dictionary (#83), more Croatian testing and fixes, microintonation for the Arató mode, a reshaped English melody, the loudness contour for UK, Australian and Canadian English, English inside Russian text (#118), and Edu's sixteen new voice profiles.
Credits
gregodejesus2 for more reports, blind rounds and regression catches than anyone; 29-Bloo for the Spanish and SAPI reports and blind rounds; edu-fblind for Brazilian Portuguese, language switching, the log that solved #127, voice profiles and pushing the editor until it made sense; MarioPercinic for Croatian and the fast-rate report; TurkrosoftLabs for Australian English, Italian and Dutch; sevapopov2 for the iOS 27 pauses and the language-switching report; Vsevolod for word-final stops; dgomez42 for the Spanish diagnosis and the Arató pitch report; LeonardoBlancoSerna, rmcpantoja and yaresDg for Spanish research and testing; kaveinthran for iOS sample-rate findings; Christopher Toth for the Klatt and DECtalk research notes; Arató András and Vaspöri Teréz, the designers of BraiLab; and GPT-6 Astra for review rounds.
Built with Claude (Anthropic) as engineering partner.
— Tamas + Claudeo