What's New
Fine-tune Whisper on your accent or domain, locally. whisper-tiny (39M) and base (74M) train on a 4 GB GPU — a new modality that actually runs on a laptop card.
task='asr'— newAsrTrainerWrapper(HFSeq2SeqTrainer+WhisperProcessor). Rows are{"audio": <path>, "text": <transcript>}underdata.format='asr'; audio decodes to 16 kHz mono through the hardened loader (soundfile pre-probe + symlink/size guards). Opt-in LoRA on q/v viatraining.asr_lora: true(default full fine-tune);training.asr_language/asr_task(transcribe|translate) set and persist the decoder prefix for inference.soup infer --task asr— transcribe an{"audio"[, "text"]}JSONL and get per-row + corpus WER/CER when references are present. Loads a full model or a LoRA adapter dir.--asr-language/--asr-task/--audio-dir.- Pure-python metrics — new
soup_cli.utils.asr_metrics(WER / CER /word_accuracy/corpus_wer), no new dependency. - Recipes —
whisper-tiny-asr/whisper-base-asr(live-trainable),whisper-large-v3-asr(parse-only),smolvlm-256m-sft. Catalog 138 → 142.
base: openai/whisper-tiny
task: asr
data:
format: asr # rows: {"audio": "clip.wav", "text": "hello world"}
audio_dir: ./data/audio
training:
asr_language: en
asr_lora: true # optional; default = full fine-tuneInstall / Upgrade
pip install -U soup-cli
Security
- Infer-path audio paths are containment-checked (realpath + commonpath) against
--audio-dir(or the input dir); traversal and Windows UNC/network paths are rejected. Per-row skip-and-warn on a bad path; atomic output write; row + character DoS caps. - Whisper base is arch-guarded (
AutoConfigmodel_type=='whisper') before any weight download;trust_remote_codeis deny-by-default. Theasr_generation.jsondecode-prefix sidecar is shape-validated on read.
Known Limitations
- WER/CER use a light normalizer, not the full Whisper English text normalizer — valid for before/after deltas, not leaderboard-comparable absolutes.
whisper-large-v3-asris parse-only (needs a larger GPU); tiny/base are live on 4 GB.- SmolVLM-256M vision recipe ships parse-tested — a live vision-SFT smoke surfaced a pre-existing Idefics3-processor incompatibility in the shared LLaVA path; Idefics3 vision-path support is a tracked follow-up.
- Validated as proof-of-mechanism (synthetic memorization: WER 1.000 → 0.000 on an RTX 3050); real accent/domain gains need real audio.
🤖 Generated with Claude Code