What's New
v0.40.6 — ReLoRA + surgical PEFT on every trainer: closes the v0.39.0 known gap. The ReLoRA callback and the Gemma 4 / fused-MoE PEFT patches now run on every transformer-backend trainer, not just SFT.
- ReLoRA × 12 trainers — set
training.relora_steps: <N>on any ofsft / dpo / grpo / kto / orpo / simpo / ipo / ppo / reward_model / pretrain / embedding / bco(or unifiedtask: preference) and the magnitude-prune-and-reset cycle fires every N steps. Schema cross-validator only rejects the MLX backend (the callback is HF Trainer-specific). - Surgical PEFT patches everywhere — Gemma 4
ClippableLinear→nn.Linearswap (so PEFT's matcher recognises the layer) and 3-D fused-MoE expert dropout strip (soParamWrapperno longer crashes on multi-expert weight tensors) now apply across every non-SFT trainer. Both patches are best-effort and architecture-gated. - One shared wiring path — new
soup_cli.utils.peft_wiringmodule exposesapply_pre_lora_patches,apply_post_lora_patches,attach_relora_callback. SFT migrated to the same helpers in the same release, so future patches land in one place and never drift between SFT and non-SFT. - +61 net new tests — source-level invariant matrix proving all 12 trainer files invoke the helpers in the right order around
get_peft_model, behavioural unit tests for each helper (Gemma 4 happy + exception swallow, post-LoRA strip happy + exception swallow, ReLoRA policy field forwarding), schema-gate matrix covering every transformer task plus thepreferencedispatcher.
Install / Upgrade
pip install --upgrade soup-cliSecurity
attach_relora_callbackusesif relora_steps is None:(project policy) so a schema-bypassing caller passingrelora_steps=0surfaces as a loudReLoRAPolicyValueError rather than a silent skip — matches v0.34.0 / v0.39.0 / v0.40.3is Noneover falsy guards.- Direct attribute access on
tcfg.relora_warmup_ratio/_reset_optimizer/_prune_ratio(nogetattrdefaults) — Pydantic schema guarantees these fields, and a misnamed attr now fails loudly withAttributeErrorinstead of silently using a wrong default. - Patch helpers swallow broad
Exceptionat DEBUG level (matches v0.39.0 best-effort design);%sformatting onexc(notrepr) so$HOME-prefixed paths cannot leak (matches v0.34.0crash.pyredaction policy). - Underlying
apply_gemma4_clippable_patchandstrip_lora_dropout_for_3d_expertsalready validatemodel_name(null bytes, length) and are duck-typed via v0.39.0 review fixes.
Known Limitations
- Real-world correctness of ReLoRA on RLHF tasks (PPO / RewardModel) is unverified — ReLoRA was originally validated on SFT/causal-LM training; rejection-sampling-style RL loops may interact unexpectedly with periodic LoRA pruning + optimizer reset. Tracked as a community QA item; the schema does not gate on this since the upstream paper does not preclude RL use.
- Multi-modal trainers (vision/audio paths in
sft.py) inherit ReLoRA + surgical patches because they share the SFT trainer wrapper, but the surgical patches are best-effort (try/except DEBUG-logged) — a Gemma 4 vision model is unlikely in practice; if encountered the patch attempt may noisy-log without applying.