github FluidInference/FluidAudio v0.17.5
v0.17.5 — Phonon-2

4 hours ago

Phonon-2 — AsrModelVersion.phonon2 (FluidInference/phonon-2-coreml). Fermion Research's quantization-aware re-training of Parakeet TDT v3 for English, in which every encoder weight takes one of five learned values per row. The Core ML build keeps those weights exactly (a sparsity mask plus fp16 palettes, iOS 18 / macOS 15 ops) — no re-quantization — and is the fastest v3-family encoder on the Neural Engine. Existing apps are unaffected: .v3 stays the default; on iOS 17 / macOS 14 .phonon2 throws and points to .ultra.

Full sets, M5 Pro, ANE, same session v3 Ultra Phonon-2
Download ~480 MB ~630 MB ~360 MB
LibriSpeech test-clean WER 2.27 % 2.13 % 2.47 %
LibriSpeech test-other WER 4.12 % 3.81 % 4.62 %
test-clean / test-other RTFx 149–152× / 138× 151× / 142× 159× / 146×
60-minute Earnings-22 file 335× 457× 480×

English only. The HF repo also carries four more exact encoders (176 MB to 470 MB) for other size/speed trade-offs; see Phonon2.md.

let models = try await AsrModels.downloadAndLoad(version: .phonon2)

CLI: --model-version phonon2 on transcribe and the benchmarks; transcribe gains --encoder-compute-units ane|gpu|cpu|all.

Also in this release: LocalVQE acoustic echo cancellation (beta, #930), Kokoro ANE fixes for long utterances (#963, #965) and English mixing in other variants (#970), LuxTTS mid-phrase pause fix (#942), an iterative word aligner for custom vocabulary (#962), and a roman-numeral list-marker fix in English text normalization (#974).

What's Changed

  • feat(asr): Phonon-2 (FermionResearch five-value v3, English) via AsrModelVersion.phonon2 by @Alex-Wengg in #980
  • feat(enhancement): LocalVQE AEC with safe streaming and benchmark validation (beta) in #930
  • fix(tts/kokoro-ane): quiet onset on long utterances — fp32 KokoroProsody_v2 in #963
  • fix(tts/kokoro-ane): restore long-text chunking in synthesizeDetailed(text:) in #965
  • feat(tts/kokoro-ane): englishPhonemes(for:) for mixing English into other variants in #970
  • fix(tts/luxtts): remove spurious mid-phrase pauses and chunk long text in #942
  • fix(vocab): make alignBaseWordsToUTF8Ranges iterative in #962
  • fix(tts/english-tn): read roman-numeral list markers as numbers in #974
  • docs: issue ↔ code ↔ HF traceability convention for model uploads in #964

Full Changelog: v0.17.4...v0.17.5

Don't miss a new FluidAudio release

NewReleases is sending notifications on new releases.