Four silent-failure modes Soup had → loud failures, plus a security default-deny.
What's New
- Assistant-only loss masking (default
true) — Soup now masks every non-assistant token withIGNORE_INDEX(-100) so the SFT loss reflects only what the model should generate. Replaces TRL's heuristic that produced wrong loss labels on intermediate user/system turns. Newdata.train_on_responses_onlyanddata.train_on_messages_with_train_fieldconfig fields. Per-messagetrain: boolopt-in for fine-grained control. Mirrors LlamaFactory + Axolotl. --trust-remote-codeopt-in (default deny) —soup train,chat,serve,data download,eval autonow refuse to load HF models that ship custom Python (auto_map) unless you pass--trust-remote-code. Allowlist of 15 first-party orgs (Meta, Mistral, Qwen, Google, etc.) suppresses warning panel. Replaces 9 unconditionaltrust_remote_code=Truecall sites in the SFT path.- Hard error on missing chat-template — Soup no longer silently builds garbage
f"{role}: {content}"when the tokenizer ships no chat template. Tokenizers without one now raise loudly with a fix suggestion. Newdata.chat_templatefield accepts a registered name (chatml/llama3/qwen2.5/mistral/gemma3/phi4/deepseek-r1) or a raw Jinja string. Filesystem-touching Jinja directives ({% include %},{% import %},{% from %},{% macro %},{% extends %}) are blocked at config load. - OOM-probe auto batch size + cache —
auto_batch_size_strategy: probe(defaultauto) replaces the static memory formula with a real try-halve-then-double-to-ceiling loop bounded by max 8 doublings, ceiling = static × 4. Picked size cached at~/.soup/batch_cache.json(0600 perms, env override containment-checked) keyed on(model, max_length, quantization, lora_r, gpu, gpu_memory_gb). - Net +134 tests (4115 → 4249) covering all four correctness fixes, Jinja directive blocklist, cache containment, trust-remote-code resolution gate, and Mistral system-turn rendering.
Install / Upgrade
pip install --upgrade soup-cli
# or, with the typical install: pip install -U "soup-cli[serve,data,eval]"Security
--trust-remote-codeopt-in: replaces 9 hardcodedtrust_remote_code=Truecall sites in the SFT path.KNOWN_SAFE_PREFIXESallowlist (15 first-party orgs) suppresses warning panel for trusted repos.chat_templatevalidator rejects null bytes, oversize (>64KB), and filesystem-touching Jinja directives._REGISTRYof named templates wrapped inMappingProxyTypeso callers cannot mutate the registry at runtime.apply_chat_template_overrideemits yellow advisory when active so users knowsoup pushwill persist the override.- Loss-mask fallback passes
add_special_tokens=Falseto incremental tokenize calls (no double-BOS drift); narrowed exception catch from(TypeError, ValueError)toTypeErroronly. SOUP_BATCH_CACHE_PATHenv override containment-checked viaos.path.realpath + commonpathagainst~/cwd/tempfile.gettempdir(). Cache file gets best-effort0o600perms (matches v0.26.0 registry.db policy).make_cache_keyrejectsboolin numeric inputs (matches v0.30.0Candidatepolicy).
Known Limitations
- Non-SFT trainers still hardcode
trust_remote_code=True: only the SFT trainer wires the new flag. DPO / GRPO / KTO / ORPO / SimPO / IPO / PPO / RewardModel / Pretrain / Embedding wrappers, pluscommands/{diff,export,merge,infer,generate}.py, still passtrust_remote_code=Trueunconditionally. Tracked for a v0.36.x patch. - Loss-mask fallback is loose: when the tokenizer doesn't support
return_assistant_tokens_mask(i.e. no{% generation %}markers in the chat template), the role-prefix tokens (e.g.<|assistant|>) end up in the loss too. Strict assistant-content-only requires a tokenizer with{% generation %}markers. - OOM probe is config-only in v0.36.0:
auto_batch_size_strategy: probereads the cache and returns the static estimate when noprobe_fnis wired (CPU-only runs, or pre-probe code path). Live CUDA probe-fn wiring is reserved for a v0.36.x patch. - License migration (MIT → Apache-2.0): applied 2026-04-21. Downstream redistributors must retain the
NOTICEfile per Apache-2.0 §4(d). Not a code follow-up — advisory only.