Automated release: version and notes generated from pull requests merged since 2.6.0.
- feat: Parakeet speech engines, built by @rymalia in #305.
parakeet-mlxandparakeet.cpp(parakeet-cli+ a.gguf) forcaption.py --transcribeandsilence.py --filler --transcribe;--engine/FFMPEG_SKILL_ASR_ENGINEpick one, and a named engine runs as asked.autoruns Parakeet for English speech (an English--language, else whisper.cpp's language detector) and Whisper otherwise. Speech whose language nobody named and nothing here could detect goes to a Whisper engine that can run; Parakeet runs on English assumed only when there is none or every one failed, and the result's top-levelnotesthen says the transcript is wrong if the speech is not English (new forsilence.py). Results carrytranscription(engine, model, language, routing) in every caption--mode;silence.py'sfiller.sourceisparakeet:ENGINEfor them. A Parakeet engine whose output is not a readable transcript is a failed run and the next engine is tried; its own empty answer is theno_speechrefusal.transcribe_result()/transcribe_words_result()return what a call used as aTranscription, andtranscribe()/transcribe_words()keep their return shapes. New optional contract capabilityexternal:parakeet. - fix: whisper.cpp is told to detect the language (
-l auto) when--language/--filler-langis not given. It was given no-l, and its own default is English, socaption.py --transcribeandsilence.py --filler --transcribedecoded speech in any other language as English. The result'stranscription.languagestaysnullfor a language the engine identified itself. - fix: the "no local speech-to-text engine found" refusal, with its install lines, now only means no engine was found. A whisper.cpp that was installed but failed, with nothing after it, got that refusal; now the refusal names each engine that was found and why it did not transcribe (
kind: input,reason: "engine_failed"or"english_only",engines[].detailwith the engine's own error line). - fix: a faster-whisper that raises while loading or running its model (an offline first run, a broken install) is a failed engine:
caption.py --transcribelogs "faster-whisper found but failed" and tries the next engine instead of refusing with "faster-whisper found no speech", andsilence.py --filler --transcribecan fall back to Parakeet. The exception used to be swallowed in the engine's thread. An engine that ran and found no segment is still theno_speechrefusal. - docs: parakeet-mlx downloads its default model (
mlx-community/parakeet-tdt-0.6b-v2) from Hugging Face on first use, as faster-whisper does, and openai-whisper downloads its model from OpenAI. The install hint, README, docs/contract.md and references/scripts.md now say so, and that the audio never leaves the machine (inference is local). The skill's own code opens no connection, but faster-whisper runs inside the skill's process, so its first-run download is made from that process. The refusal's "all run offline" is gone: a first run that downloads is not offline. - feat(contract):
execution.model_downloadssays which speech engine fetches its model on first use, from where, and in which process (faster-whisper in the skill's own; parakeet-mlx and openai-whisper in their own child processes; whisper.cpp and parakeet.cpp never).execution.network: falsekeeps its meaning, the skill's own code, and no longer reads as a promise that a first--transcribeis network-free. - feat: Parakeet speech engines for --transcribe (#312)
What's Changed
Full Changelog: v2.6.0...v2.7.0