What's New
Seven opt-in flags that turn Soup into the "fast.ai of LLM fine-tuning" — pre-flight LR tuning + in-training stability nets.
- LR Range Finder —
soup train --find-lrruns a fast.ai-style geometric LR sweep and writes a JSON report with the recommended LR, the EMA-smoothed loss curve, and the divergence point. - Auto warmup schedule —
training.warmup_auto: truederiveswarmup_stepsfromdataset_size × epochs × warmup_ratio, clamped to[10, 1000].warmup_ratio: 0.0short-circuits to "no warmup" matching HF Trainer's convention. - Auto mixed-precision —
training.auto_mixed_precision: truepicksbf16(Ampere+),fp16(Turing or known fp16-stable models like Qwen2 / Phi-3.5), orno(Pascal and older). Multi-version pairs (qwen2.5vsqwen2,phi-3.5vsphi-3) match by longest substring. - Loss spike auto-recovery — extends the watchdog: when loss spikes, decay LR and resume instead of dying.
loss_spike_recovery: trueon top ofloss_watchdog: true. Capped at 3 attempts by default. - Convergence detector —
convergence_detection: truesurfacescontinue/early_stop/lower_lradvice from the loss curve. - VRAM-pressure advisory —
grad_accum_auto_tune: truerecords peak memory each step and recommends a new(batch, accum)pair (capped ataccum=1024) when pressure crosses your threshold. - Autopilot integration — the zero-config decision engine now picks
warmup_stepsand mixed-precision automatically when you runsoup autopilot.
Install / Upgrade
```bash
pip install --upgrade soup-cli
or
docker pull ghcr.io/makazhanalpamys/soup:v0.32.0
```
Security
14 new hardening lines documented in SECURITY.md: `--find-lr-output` containment via shared `is_under_cwd`; NaN/Infinity rejection + `allow_nan=False` JSON; `compute_lr_schedule` rejects non-positive / inverted ranges + `num_steps` outside [2, 10_000]; `pick_mixed_precision` rejects empty / null-byte / >200-char names; longest-substring quirk lookup; `compute_warmup_steps` clamps to [10, 1000]; `SpikeRecoveryStrategy` is `@dataclass(frozen=True)`; cross-validator rejects `loss_spike_recovery` without `loss_watchdog`; `convergence_*` bounded; `GradAccumMonitor.recommend()` caps doubled accum at `MAX_ACCUM=1024`; `generate_config` validates BOTH the YAML output path AND embedded `decisions["output"]`.
Known Limitations
The schemas, validators, picker APIs, and JSON report format ship in v0.32.0 and are stable. Live in-process wiring lands in v0.32.1:
- `--find-lr` writes a stub report from a synthetic descend-then-explode curve. The real in-process LR-sweep training loop lands in v0.32.1. The flag prints a yellow advisory so it is never a silent no-op (same pattern as v0.30.0 `--auto-quant`).
- `loss_spike_recovery` ships the policy + cross-validator. Live trainer-state rollback + LR decay + resume wiring lands in v0.32.1.
- `grad_accum_auto_tune` is advisory only — recommends a new (batch, accum) pair but does not mutate the live DataLoader. Live mutation requires TRL changes.
- `auto_mixed_precision` ships the picker API + config field. SFTTrainerWrapper integration to push the picked precision into HF `TrainingArguments` lands in v0.32.1.
Stats
- 89 new tests (3607 → 3696 passing)
- ruff clean
- Five-agent review wave: every CRITICAL / HIGH / MEDIUM / LOW finding fixed before tag
🤖 Generated with Claude Code