github MakazhanAlpamys/Soup v0.71.29
v0.71.29 — soup shrink: depth-prune + distill-heal

latest releases: v0.75.2, v0.75.1, v0.75.0...
3 months ago

What's New

soup shrink — make your model smaller, locally. Depth-prune the least-important contiguous block of decoder layers (ranked by the angular distance of the residual stream across each block over a calibration set — "The Unreasonable Ineffectiveness of the Deeper Layers", arXiv:2403.17887), optionally distill-heal the loss (Minitron-style logit KD), and get a binary SHIP / DON'T-SHIP perplexity verdict. No other fine-tuning CLI ships this.

  • Importance-ranked depth pruning. One output_hidden_states forward per calib prompt scores every candidate block; the least-important one is dropped (first + last layer always protected).
  • Distill-heal to one dense model. --heal distills the full-depth original into the pruned student (LoRA logit-KD) as an isolated soup train run, then fuses the adapter back — you ship a single dense smaller model, not base + adapter.
  • A verdict, not a dashboard. exit 0 = SHIP, 2 = DON'T SHIP, 1 = error — drops straight into CI. --plan-only prints the importance table without writing anything.
  • Runs on a laptop GPU. Live-validated on Windows + RTX 3050 with SmolLM2-135M: drop 25% (30→22 layers, 21% params, ppl ×2.98) and drop-4 + CPU heal (ppl recovered to ×1.35).
soup shrink --model HuggingFaceTB/SmolLM2-135M-Instruct --drop-ratio 0.25 \
    --calib calib.jsonl --heal heal.jsonl --heal-steps 200 -o shrunk --device cpu

Install / Upgrade

pip install -U soup-cli

Security

  • soup shrink contains --calib / --heal / --output-dir and every derived write path (<out>/model, <out>/heal_adapter, the fuse staging dir) under cwd via realpath + commonpath + O_NOFOLLOW + symlink rejection, re-validated right before each write (TOCTOU defence over the potentially-long heal).
  • The heal subprocess uses an argv list (no shell) with a timeout; its config is schema-validated before spawn; subprocess output is C0/ESC-stripped before it reaches the terminal. --model defaults trust_remote_code=False with a probe + warn.

Known Limitations

  1. The importance pass loads the full model → live-validated ≤ 3 B on the 4 GB reference box (SmolLM2-135M); larger models work but are unvalidated on that hardware.
  2. Heal keeps the teacher resident (~2× model memory); the LoRA student + --device cpu fallback (CUDA_VISIBLE_DEVICES=-1) keep it tractable, but a real GPU heal of > ~1.5 B on 4 GB hits the hardware-fit gate — use --device cpu or a bigger card.
  3. Arch allowlist v1 = Llama / Qwen / SmolLM; other families are a friendly reject (a follow-up tracks widening it).
  4. Perplexity is an unweighted mean of per-example perplexities — valid for the before/after ratio the verdict uses, not directly comparable to soup eval / lm-eval absolute numbers.
  5. --heal-steps maps to epochs (no max_steps knob) via ceil(steps*batch/rows), clamped to 100 epochs.

Full changelog: CHANGELOG.md

Don't miss a new Soup release

NewReleases is sending notifications on new releases.