What's New
Patch release closing 14 bugs from the v0.40.0 QA pass on Windows + RTX 3050 4 GB, plus the deferred multi-objective preference loss runtime stub.
- UTF-8 stdio bootstrap (Windows) — one-line fix at the CLI entrypoint reconfigures stdout/stderr to UTF-8 before any Rich console init. Closes 6 separate QA findings (β / ✓ / box-drawing crashes on cp1251/cp1252). POSIX no-op.
- Multi-objective preference loss is live —
training.preference_loss_weights: {dpo: 0.7, simpo: 0.3}no longer raisesNotImplementedError. Builds the highest-weighted loss as primary; rejects BCO + paired loss blends at runtime (data-format incompatible). Pure-function math kernel exercised by 24 tests. - Smarter defaults on small GPUs —
soup quickstartauto-switches toSmolLM2-135M-Instructon ≤6 GB VRAM (verified in 5 s on RTX 3050 4 GB); autopilot fallback now defaults to 1B instead of 7B and reads the local safetensors index when available. soup doctorupgrades — flagstransformers ≥ 5.0.0as INCOMPATIBLE; distinguishes "no GPU hardware" from "GPU hardware present, wrong torch wheel" (nvidia-smisucceeds buttorch.cuda.is_available()returns False); detects dual-Python interpreter setups; usesimportlib.metadataas version-probe fallback.- CLI UX consistency —
soup init --forcenon-interactively overwrites;--templatehelp generated from the live registry;soup migrate <data.jsonl>errors with "did you pass the wrong file?";soup eval custom -onow writes JSON regardless of--attach-to-registry;soup recipes show <typo>suggests close matches viadifflib;soup data samplefilenames embed strategy (no overwrite); JSONL loader auto-strips UTF-8 BOM. - Schema strictness — root-level
lora:in YAML (the LlamaFactory / Axolotl convention) now migrates intotraining.loraso nested validators (includinginit_strategy: random|pissa|olora) actually fire instead of being silently dropped. --find-lractually runs the live loop — fixed brokenload_localimport that previously always silently fell through to a static placeholder curve.
Install / Upgrade
pip install -U soup-cliOr via Docker (GHCR):
docker pull ghcr.io/makazhanalpamys/soup:v0.40.1Security
- UTF-8 stdio bootstrap swallows
(OSError, ValueError, AttributeError)on detached streams; never overrides user-setPYTHONIOENCODING. _remap_root_level_misplaced_keysdoes not mutate caller dict (shallow-copy policy mirroring v0.33.0 #47 / v0.40.0 Part B)._probe_cache_param_countrejects empty / null-byte model names before path construction.nvidia-smiGPU label isrich.markup.escaped before embedding in markup string (NVIDIA Quadro [T4]cannot break or inject markup).- Dual-Python detector uses
os.path.realpath(notPath.resolve()) for Windows 8.3 short-name compat. combine_lossesrejects empty weights, propagates NaN loudly, rejectsboolweight values.
Known Limitations
- Multi-objective preference loss is a primary-loss approximation — the highest-weighted loss is selected as the primary inner trainer; auxiliary losses are named in the advisory but do not yet contribute to the backward pass. Full per-batch weighted-loss combination across all named losses (one forward pass shared across DPO/SimPO/ORPO/IPO via TRL trainer subclassing) is deferred to v0.40.2.
--trust-remote-codeopt-in surface unchanged — non-SFT trainers (DPO / GRPO / KTO / ORPO / SimPO / IPO / PPO / RewardModel / Pretrain / Embedding / BCO) pluscommands/{diff,export,merge,infer,generate}.pystill hardcodetrust_remote_code=True. Tracked under issue #63 for incremental expansion.- CI / docs / UX papercuts deferred to v0.40.2 — H2 (data flag aliases), H3 (
quickstart --output), N7 (infer/bench HF id auto-download), M4 (data dedup --threshold), M5 (runs --cwd-only), plus the post-v0.40.1 patch-chain items #36 / #50 / #51.
Test count
4720 passed / 3 skipped (was 4656 in v0.40.0 → +64 net new).