github MakazhanAlpamys/Soup v0.71.25
v0.71.25 — soup ship: the SHIP / DON'T-SHIP verdict

latest releases: v0.75.2, v0.75.1, v0.75.0...
3 months ago

What's New

soup ship — answer one question after fine-tuning: did the model get better, or did I break it? A single binary SHIP / DON'T SHIP verdict + a one-screen reason — not a dashboard to guess from.

soup ship --base <model> --adapter <lora> --task-eval tasks.jsonl   # → SHIP / DON'T SHIP
  • One decision, two legs. SHIP only when (1) the task metric strictly improved (base → tuned) AND (2) no general benchmark regressed past a forgetting threshold (default 0.05 absolute points). Otherwise DON'T SHIP — even if the task metric looks great.
  • The moat is leg 2 — a catastrophic-forgetting / regression gate, first-class and fused with the task win.
  • CI-gateable: exit 0 = SHIP, 2 = DON'T SHIP, 1 = runtime error.
  • Leg-1 modes: --task-mode metric (eval accuracy) or judge_score (LLM-as-a-judge).
  • Leg-2 suite: built-in mini benchmarks by default (offline, CPU); --general-suite <names> routes lm-eval benchmarks; --baseline registry://… | file.json supplies recorded base scores.
  • Offline: soup ship --evidence ev.json decides from pre-computed scores (no model load); --output verdict.json persists the machine-readable verdict.

Validated live on a real SmolLM2-135M base vs a LoRA tune (SHIP, exit 0). Engine (soup_cli.utils.ship_verdict.decide_ship) is pure-python and CPU-testable.

Install / Upgrade

pip install -U soup-cli            # light CLI
pip install -U 'soup-cli[train]'   # + training stack

Security

soup ship input hardening: --evidence opened with O_NOFOLLOW + an fstat size cap (16 MiB) under cwd containment; --task-eval cwd-contained and symlink-rejected; --judge-model validated by scheme/host via urlparse (blocks the http://localhost.attacker.com prefix bypass); lm-eval model ids reject ,/= injection (trust_remote_code smuggling); --general-suite bounded (≤ 50 names, ≤ 256 chars each); leg 2 refuses (exit 1) rather than silently shipping when a requested benchmark cannot be scored on both sides.

Known Limitations

  • Pairwise judge win-rate not shipped — --task-mode pairwise is in the engine enum but rejected by the CLI; leg-1 v1 is metric + judge_score. Planned fast-follow.
  • ShipConfig schema deferred — soup ship is CLI-only in v1 (no soup.yaml block).
  • Pre-existing judge-URL bypass in eval/gate.py — the same startswith("http://localhost") prefix bypass exists in the v0.26.0 gate judge-URL allowlist (soup eval gate / soup train --gate); soup ship ships its own urlparse validator. Back-port tracked as a follow-up issue.

Don't miss a new Soup release

NewReleases is sending notifications on new releases.