github MakazhanAlpamys/Soup v0.71.32
v0.71.32 — ASR fine-tuning (Whisper)

latest releases: v0.75.2, v0.75.1, v0.75.0...
3 months ago

What's New

Fine-tune Whisper on your accent or domain, locally. whisper-tiny (39M) and base (74M) train on a 4 GB GPU — a new modality that actually runs on a laptop card.

  • task='asr' — new AsrTrainerWrapper (HF Seq2SeqTrainer + WhisperProcessor). Rows are {"audio": <path>, "text": <transcript>} under data.format='asr'; audio decodes to 16 kHz mono through the hardened loader (soundfile pre-probe + symlink/size guards). Opt-in LoRA on q/v via training.asr_lora: true (default full fine-tune); training.asr_language / asr_task (transcribe|translate) set and persist the decoder prefix for inference.
  • soup infer --task asr — transcribe an {"audio"[, "text"]} JSONL and get per-row + corpus WER/CER when references are present. Loads a full model or a LoRA adapter dir. --asr-language / --asr-task / --audio-dir.
  • Pure-python metrics — new soup_cli.utils.asr_metrics (WER / CER / word_accuracy / corpus_wer), no new dependency.
  • Recipes — whisper-tiny-asr / whisper-base-asr (live-trainable), whisper-large-v3-asr (parse-only), smolvlm-256m-sft. Catalog 138 → 142.
base: openai/whisper-tiny
task: asr
data:
  format: asr            # rows: {"audio": "clip.wav", "text": "hello world"}
  audio_dir: ./data/audio
training:
  asr_language: en
  asr_lora: true         # optional; default = full fine-tune

Install / Upgrade

pip install -U soup-cli

Security

  • Infer-path audio paths are containment-checked (realpath + commonpath) against --audio-dir (or the input dir); traversal and Windows UNC/network paths are rejected. Per-row skip-and-warn on a bad path; atomic output write; row + character DoS caps.
  • Whisper base is arch-guarded (AutoConfig model_type=='whisper') before any weight download; trust_remote_code is deny-by-default. The asr_generation.json decode-prefix sidecar is shape-validated on read.

Known Limitations

  1. WER/CER use a light normalizer, not the full Whisper English text normalizer — valid for before/after deltas, not leaderboard-comparable absolutes.
  2. whisper-large-v3-asr is parse-only (needs a larger GPU); tiny/base are live on 4 GB.
  3. SmolVLM-256M vision recipe ships parse-tested — a live vision-SFT smoke surfaced a pre-existing Idefics3-processor incompatibility in the shared LLaVA path; Idefics3 vision-path support is a tracked follow-up.
  4. Validated as proof-of-mechanism (synthetic memorization: WER 1.000 → 0.000 on an RTX 3050); real accent/domain gains need real audio.

🤖 Generated with Claude Code

Don't miss a new Soup release

NewReleases is sending notifications on new releases.