What's New
soup shrink — make your model smaller, locally. Depth-prune the least-important contiguous block of decoder layers (ranked by the angular distance of the residual stream across each block over a calibration set — "The Unreasonable Ineffectiveness of the Deeper Layers", arXiv:2403.17887), optionally distill-heal the loss (Minitron-style logit KD), and get a binary SHIP / DON'T-SHIP perplexity verdict. No other fine-tuning CLI ships this.
- Importance-ranked depth pruning. One
output_hidden_statesforward per calib prompt scores every candidate block; the least-important one is dropped (first + last layer always protected). - Distill-heal to one dense model.
--healdistills the full-depth original into the pruned student (LoRA logit-KD) as an isolatedsoup trainrun, then fuses the adapter back — you ship a single dense smaller model, not base + adapter. - A verdict, not a dashboard.
exit 0 = SHIP,2 = DON'T SHIP,1 = error— drops straight into CI.--plan-onlyprints the importance table without writing anything. - Runs on a laptop GPU. Live-validated on Windows + RTX 3050 with SmolLM2-135M: drop 25% (30→22 layers, 21% params, ppl ×2.98) and drop-4 + CPU heal (ppl recovered to ×1.35).
soup shrink --model HuggingFaceTB/SmolLM2-135M-Instruct --drop-ratio 0.25 \
--calib calib.jsonl --heal heal.jsonl --heal-steps 200 -o shrunk --device cpuInstall / Upgrade
pip install -U soup-cliSecurity
soup shrinkcontains--calib/--heal/--output-dirand every derived write path (<out>/model,<out>/heal_adapter, the fuse staging dir) under cwd viarealpath+commonpath+O_NOFOLLOW+ symlink rejection, re-validated right before each write (TOCTOU defence over the potentially-long heal).- The heal subprocess uses an argv list (no shell) with a timeout; its config is schema-validated before spawn; subprocess output is C0/ESC-stripped before it reaches the terminal.
--modeldefaultstrust_remote_code=Falsewith a probe + warn.
Known Limitations
- The importance pass loads the full model → live-validated ≤ 3 B on the 4 GB reference box (SmolLM2-135M); larger models work but are unvalidated on that hardware.
- Heal keeps the teacher resident (~2× model memory); the LoRA student +
--device cpufallback (CUDA_VISIBLE_DEVICES=-1) keep it tractable, but a real GPU heal of > ~1.5 B on 4 GB hits the hardware-fit gate — use--device cpuor a bigger card. - Arch allowlist v1 = Llama / Qwen / SmolLM; other families are a friendly reject (a follow-up tracks widening it).
- Perplexity is an unweighted mean of per-example perplexities — valid for the before/after ratio the verdict uses, not directly comparable to
soup eval/ lm-eval absolute numbers. --heal-stepsmaps to epochs (nomax_stepsknob) viaceil(steps*batch/rows), clamped to 100 epochs.
Full changelog: CHANGELOG.md