github FluidInference/FluidAudio v0.13.7

latest releases: v0.17.7, v0.17.6, v0.17.5...
5 months ago

What's New in v0.13.7

Features

  • Parallelize chunked Parakeet batch transcription (#507) — ~2.2–2.8× speedup on long audio via ASRConfig.parallelChunkConcurrency, no change to streaming/live path.
  • Nemotron 160 ms and 80 ms chunk sizes (#490) — new NemotronChunkSize.ms160 / .ms80, plus Repo.nemotronStreaming160 / nemotronStreaming80.
  • Japanese TDT via unified AsrModels API (#521) — Japanese TDT models now work through AsrModels / AsrManager with timing info; redundant TdtJaManager / CtcJaManager removed.
  • Custom segment activity reporting for diarization (#493) — DiarizerActivityType.sigmoids (default, existing behaviour) or .logits.
  • Lower ASR minimum-audio guard from 1 s to 300 ms (#531) — short single-word utterances ("yes", "no", "stop") no longer rejected before inference. Adds ASRConstants.minimumAudioDurationSeconds.

Fixes

  • Offline diarization single-speaker output (#523) — three bugs: vDSP_mtrans dimension swap, missing <20% activity-ratio filter, soft vs binary per-frame masks. Now matches the reference pyannote pipeline.
  • Japanese TDT infinite re-download loop (#522) — download() now uses version-specific filenames (Decoderv2.mlmodelc, Jointerv2.mlmodelc) so modelsExist() stops failing the existence check.
  • parakeet-ctc-ja load error (#516) — AsrModels no longer accepts .ctcJa / .ctcZhCn (which use CtcDecoder.mlmodelc, not the TDT Decoder.mlmodelc).
  • Benchmark script / download bugs (#534) — repair three pre-existing bugs blocking Scripts/parakeet_subset_benchmark.sh --download.

Refactoring

  • Standardized model loading API across all ASR managers (#506) — unified loadModels entry points replacing ad-hoc configure / startStreaming / loadModels(modelDir:) mix.
  • Reorganize batch managers + expose decoder state explicitly (#502) — batch managers grouped under SlidingWindow/ by algorithm (TDT vs CTC); decoder state no longer hidden behind per-source routing.
  • Deduplicate language-specific model files (#492) — ~700 lines of boilerplate folded into a generic ParakeetLanguageModels.
  • Consistency pass across Parakeet ASR managers (#494) — standardized lifecycle method names and related cleanups.
  • Rename Repo.parakeetCtcJa → Repo.parakeetJa (#520) — the repo contains both CTC and TDT weights; the old name was misleading.
  • AudioConverter is now Sendable (#505).

Documentation

  • Complete API reference and updated ASR documentation (#498)
  • Reorganized Documentation/README.md for discoverability (#497)
  • Refined model descriptions in Models.md (#496, #501)
  • Diarization docs: clarify pipeline version differences (#511) and fix 3.1 → community-1 model references (#510)
  • Update ModelConversion.md with PR referencing / validation steps (637b609)
  • Update AMI offline benchmarks after the #523 fix (#533)
  • Add git-worktree guidance for multi-agent workflows (#535)

CI / Chores

  • Simplify TTS smoke-test PR comments (#504)
  • Remove personal VSCode settings (#500) and orphaned arm64-build.png (#499)

New Contributors

  • @hamzaq2000 made their first contribution in #507 (Parallelize chunked Parakeet batch transcription)
  • @thechatk made their first contribution in #523 (Fix offline diarization pipeline producing single-speaker output)
  • @esphoenixc made their first contribution in #531 (Lower ASR minimum audio guard from 1 s to 300 ms)

Full Changelog: v0.13.6...v0.13.7

Don't miss a new FluidAudio release

NewReleases is sending notifications on new releases.