github MakazhanAlpamys/Soup v0.40.6
v0.40.6 — ReLoRA + surgical PEFT on every trainer

latest releases: v0.75.2, v0.75.1, v0.75.0...
4 months ago

What's New

v0.40.6 — ReLoRA + surgical PEFT on every trainer: closes the v0.39.0 known gap. The ReLoRA callback and the Gemma 4 / fused-MoE PEFT patches now run on every transformer-backend trainer, not just SFT.

  • ReLoRA × 12 trainers — set training.relora_steps: <N> on any of sft / dpo / grpo / kto / orpo / simpo / ipo / ppo / reward_model / pretrain / embedding / bco (or unified task: preference) and the magnitude-prune-and-reset cycle fires every N steps. Schema cross-validator only rejects the MLX backend (the callback is HF Trainer-specific).
  • Surgical PEFT patches everywhere — Gemma 4 ClippableLinear → nn.Linear swap (so PEFT's matcher recognises the layer) and 3-D fused-MoE expert dropout strip (so ParamWrapper no longer crashes on multi-expert weight tensors) now apply across every non-SFT trainer. Both patches are best-effort and architecture-gated.
  • One shared wiring path — new soup_cli.utils.peft_wiring module exposes apply_pre_lora_patches, apply_post_lora_patches, attach_relora_callback. SFT migrated to the same helpers in the same release, so future patches land in one place and never drift between SFT and non-SFT.
  • +61 net new tests — source-level invariant matrix proving all 12 trainer files invoke the helpers in the right order around get_peft_model, behavioural unit tests for each helper (Gemma 4 happy + exception swallow, post-LoRA strip happy + exception swallow, ReLoRA policy field forwarding), schema-gate matrix covering every transformer task plus the preference dispatcher.

Install / Upgrade

pip install --upgrade soup-cli

Security

  • attach_relora_callback uses if relora_steps is None: (project policy) so a schema-bypassing caller passing relora_steps=0 surfaces as a loud ReLoRAPolicy ValueError rather than a silent skip — matches v0.34.0 / v0.39.0 / v0.40.3 is None over falsy guards.
  • Direct attribute access on tcfg.relora_warmup_ratio / _reset_optimizer / _prune_ratio (no getattr defaults) — Pydantic schema guarantees these fields, and a misnamed attr now fails loudly with AttributeError instead of silently using a wrong default.
  • Patch helpers swallow broad Exception at DEBUG level (matches v0.39.0 best-effort design); %s formatting on exc (not repr) so $HOME-prefixed paths cannot leak (matches v0.34.0 crash.py redaction policy).
  • Underlying apply_gemma4_clippable_patch and strip_lora_dropout_for_3d_experts already validate model_name (null bytes, length) and are duck-typed via v0.39.0 review fixes.

Known Limitations

  • Real-world correctness of ReLoRA on RLHF tasks (PPO / RewardModel) is unverified — ReLoRA was originally validated on SFT/causal-LM training; rejection-sampling-style RL loops may interact unexpectedly with periodic LoRA pruning + optimizer reset. Tracked as a community QA item; the schema does not gate on this since the upstream paper does not preclude RL use.
  • Multi-modal trainers (vision/audio paths in sft.py) inherit ReLoRA + surgical patches because they share the SFT trainer wrapper, but the surgical patches are best-effort (try/except DEBUG-logged) — a Gemma 4 vision model is unlikely in practice; if encountered the patch attempt may noisy-log without applying.

Don't miss a new Soup release

NewReleases is sending notifications on new releases.