What's New in v0.13.7
Features
- Parallelize chunked Parakeet batch transcription (#507) — ~2.2–2.8× speedup on long audio via
ASRConfig.parallelChunkConcurrency, no change to streaming/live path. - Nemotron 160 ms and 80 ms chunk sizes (#490) — new
NemotronChunkSize.ms160/.ms80, plusRepo.nemotronStreaming160/nemotronStreaming80. - Japanese TDT via unified AsrModels API (#521) — Japanese TDT models now work through
AsrModels/AsrManagerwith timing info; redundantTdtJaManager/CtcJaManagerremoved. - Custom segment activity reporting for diarization (#493) —
DiarizerActivityType.sigmoids(default, existing behaviour) or.logits. - Lower ASR minimum-audio guard from 1 s to 300 ms (#531) — short single-word utterances ("yes", "no", "stop") no longer rejected before inference. Adds
ASRConstants.minimumAudioDurationSeconds.
Fixes
- Offline diarization single-speaker output (#523) — three bugs:
vDSP_mtransdimension swap, missing <20% activity-ratio filter, soft vs binary per-frame masks. Now matches the reference pyannote pipeline. - Japanese TDT infinite re-download loop (#522) —
download()now uses version-specific filenames (Decoderv2.mlmodelc,Jointerv2.mlmodelc) somodelsExist()stops failing the existence check. parakeet-ctc-jaload error (#516) —AsrModelsno longer accepts.ctcJa/.ctcZhCn(which useCtcDecoder.mlmodelc, not the TDTDecoder.mlmodelc).- Benchmark script / download bugs (#534) — repair three pre-existing bugs blocking
Scripts/parakeet_subset_benchmark.sh --download.
Refactoring
- Standardized model loading API across all ASR managers (#506) — unified
loadModelsentry points replacing ad-hocconfigure/startStreaming/loadModels(modelDir:)mix. - Reorganize batch managers + expose decoder state explicitly (#502) — batch managers grouped under
SlidingWindow/by algorithm (TDT vs CTC); decoder state no longer hidden behind per-source routing. - Deduplicate language-specific model files (#492) — ~700 lines of boilerplate folded into a generic
ParakeetLanguageModels. - Consistency pass across Parakeet ASR managers (#494) — standardized lifecycle method names and related cleanups.
- Rename
Repo.parakeetCtcJa→Repo.parakeetJa(#520) — the repo contains both CTC and TDT weights; the old name was misleading. AudioConverteris nowSendable(#505).
Documentation
- Complete API reference and updated ASR documentation (#498)
- Reorganized
Documentation/README.mdfor discoverability (#497) - Refined model descriptions in
Models.md(#496, #501) - Diarization docs: clarify pipeline version differences (#511) and fix 3.1 → community-1 model references (#510)
- Update
ModelConversion.mdwith PR referencing / validation steps (637b609) - Update AMI offline benchmarks after the #523 fix (#533)
- Add git-worktree guidance for multi-agent workflows (#535)
CI / Chores
- Simplify TTS smoke-test PR comments (#504)
- Remove personal VSCode settings (#500) and orphaned
arm64-build.png(#499)
New Contributors
- @hamzaq2000 made their first contribution in #507 (Parallelize chunked Parakeet batch transcription)
- @thechatk made their first contribution in #523 (Fix offline diarization pipeline producing single-speaker output)
- @esphoenixc made their first contribution in #531 (Lower ASR minimum audio guard from 1 s to 300 ms)
Full Changelog: v0.13.6...v0.13.7