github MakazhanAlpamys/Soup v0.40.4
v0.40.4 — trust_remote_code multi-trainer + multipack live (#63, #65)

latest releases: v0.75.2, v0.75.1, v0.75.0...
4 months ago

What's New

v0.40.4 closes two carry-over gaps from earlier releases. No new user-facing features — instead, --trust-remote-code becomes uniform across every trainer + command (closing the v0.36.0 #63 known gap), and the multipack FFD bin-packing sampler finally lands in HF Trainer's DataLoader (closing the v0.40.3-deferred #65).

--trust-remote-code everywhere

Every command that loads a model now defaults to trust_remote_code=False and only enables custom-code execution on explicit opt-in:

  • Trainers — DPO, GRPO, KTO, ORPO, SimPO, IPO, PPO, RewardModel, Pretrain, Embedding, BCO, and the unified Preference dispatcher
  • Commands — soup diff, soup export, soup merge, soup infer, soup data generate

Unknown-org local checkpoints with auto_map raise a friendly ValueError at construction time instead of silently exec'ing inside from_pretrained. The v0.36.0 KNOWN_SAFE_PREFIXES allowlist (Meta, Mistral, Qwen, Google, etc.) still suppresses the warning panel for first-party orgs.

soup train --config soup.yaml --trust-remote-code
soup infer --model my-org/custom-arch --input prompts.jsonl --trust-remote-code
soup export --model ./adapter --format gguf --trust-remote-code

Multipack — live HF Trainer wiring landed

make_multipack_trainer_class now adds a get_train_dataloader override that installs MultipackBatchSampler(real_batches=False) (yields a flat list[int] per packed sequence — DataLoader-compatible) as the DataLoader's batch_sampler=, forwarding dataloader_drop_last / num_workers / pin_memory from TrainingArguments. SFT and Pretrain trainer wrappers now actually instantiate the multipack subclass when multipack: true (the v0.40.3 yellow advisory + standard-sampler fallback is gone).

training:
  multipack: true
  packing: false   # mutually exclusive

Architecture allowlist (18 supported families) still gates at config-load. Multipack is sft / pretrain only on the transformers backend; preference / RLHF trainers and MLX backend get distinct error messages.

Install / Upgrade

pip install --upgrade soup-cli

Security

  • Closes the v0.36.0 #63 known gap: every non-SFT trainer wrapper + 5 standalone commands now thread trust_remote_code through the v0.36.0 resolve_trust_remote_code helper. No remaining trust_remote_code=True literal anywhere in soup_cli/trainer/ (asserted by tests/test_v0404_part_a.py).
  • Code-review fix on the multipack subclass: _get_train_sampler override always delegates to super() (returning a multipack list[list[int]] from this method would be a shape mismatch if any HF eval / prediction loop bypassed get_train_dataloader).
  • _load_reward_model (module-level helper in ppo.py) is now safe to call from outside PPOTrainerWrapper — accepts and resolves trust_remote_code independently.

Known Limitations

  • multipack: true requires the dataset to expose input_ids (preferred) or length per row. Un-tokenized text-only datasets trigger the v0.40.3 all-zeros WARNING; the MultipackBatchSampler will reject the run loudly. Pre-tokenization is the user's responsibility.
  • The DataLoader override does NOT thread FSDP / DeepSpeed parallelism env hints. Distributed multipack: true runs are still untested under FSDP / ZeRO — tracked for v0.40.5+ (paired with v0.42.0 multi-GPU work).
  • _live_lr_sweep_from_config in commands/train.py still hardcodes trust_remote_code=False for the LR sweep's internal model load — defensive but means --find-lr cannot consume custom-code models even with the user opt-in.
  • Each non-SFT trainer's __init__ repeats the resolver block (10 sites). Code-quality refactor candidate (single shared helper) deferred to a future patch.
  • Custom HF Space templates from v0.40.2 still create with space_sdk="gradio" — design choice carried over from v0.40.3.

Test count

4855 → 4930 (+75 net new tests across the trainer×trust_remote_code matrix, source-level invariants, the new DataLoader override, and the real transformers.Trainer MRO mix-in).

Don't miss a new Soup release

NewReleases is sending notifications on new releases.