github MakazhanAlpamys/Soup v0.32.0
v0.32.0 — Training Stability & Auto-Tuning

latest releases: v0.75.2, v0.75.1, v0.75.0...
5 months ago

What's New

Seven opt-in flags that turn Soup into the "fast.ai of LLM fine-tuning" — pre-flight LR tuning + in-training stability nets.

  • LR Range Finder — soup train --find-lr runs a fast.ai-style geometric LR sweep and writes a JSON report with the recommended LR, the EMA-smoothed loss curve, and the divergence point.
  • Auto warmup schedule — training.warmup_auto: true derives warmup_steps from dataset_size × epochs × warmup_ratio, clamped to [10, 1000]. warmup_ratio: 0.0 short-circuits to "no warmup" matching HF Trainer's convention.
  • Auto mixed-precision — training.auto_mixed_precision: true picks bf16 (Ampere+), fp16 (Turing or known fp16-stable models like Qwen2 / Phi-3.5), or no (Pascal and older). Multi-version pairs (qwen2.5 vs qwen2, phi-3.5 vs phi-3) match by longest substring.
  • Loss spike auto-recovery — extends the watchdog: when loss spikes, decay LR and resume instead of dying. loss_spike_recovery: true on top of loss_watchdog: true. Capped at 3 attempts by default.
  • Convergence detector — convergence_detection: true surfaces continue / early_stop / lower_lr advice from the loss curve.
  • VRAM-pressure advisory — grad_accum_auto_tune: true records peak memory each step and recommends a new (batch, accum) pair (capped at accum=1024) when pressure crosses your threshold.
  • Autopilot integration — the zero-config decision engine now picks warmup_steps and mixed-precision automatically when you run soup autopilot.

Install / Upgrade

```bash
pip install --upgrade soup-cli

or

docker pull ghcr.io/makazhanalpamys/soup:v0.32.0
```

Security

14 new hardening lines documented in SECURITY.md: `--find-lr-output` containment via shared `is_under_cwd`; NaN/Infinity rejection + `allow_nan=False` JSON; `compute_lr_schedule` rejects non-positive / inverted ranges + `num_steps` outside [2, 10_000]; `pick_mixed_precision` rejects empty / null-byte / >200-char names; longest-substring quirk lookup; `compute_warmup_steps` clamps to [10, 1000]; `SpikeRecoveryStrategy` is `@dataclass(frozen=True)`; cross-validator rejects `loss_spike_recovery` without `loss_watchdog`; `convergence_*` bounded; `GradAccumMonitor.recommend()` caps doubled accum at `MAX_ACCUM=1024`; `generate_config` validates BOTH the YAML output path AND embedded `decisions["output"]`.

Known Limitations

The schemas, validators, picker APIs, and JSON report format ship in v0.32.0 and are stable. Live in-process wiring lands in v0.32.1:

  • `--find-lr` writes a stub report from a synthetic descend-then-explode curve. The real in-process LR-sweep training loop lands in v0.32.1. The flag prints a yellow advisory so it is never a silent no-op (same pattern as v0.30.0 `--auto-quant`).
  • `loss_spike_recovery` ships the policy + cross-validator. Live trainer-state rollback + LR decay + resume wiring lands in v0.32.1.
  • `grad_accum_auto_tune` is advisory only — recommends a new (batch, accum) pair but does not mutate the live DataLoader. Live mutation requires TRL changes.
  • `auto_mixed_precision` ships the picker API + config field. SFTTrainerWrapper integration to push the picked precision into HF `TrainingArguments` lands in v0.32.1.

Stats

  • 89 new tests (3607 → 3696 passing)
  • ruff clean
  • Five-agent review wave: every CRITICAL / HIGH / MEDIUM / LOW finding fixed before tag

🤖 Generated with Claude Code

Don't miss a new Soup release

NewReleases is sending notifications on new releases.