github MakazhanAlpamys/Soup v0.71.26
v0.71.26 — Closed-loop reward-hacking auto-mitigation

latest releases: v0.75.2, v0.75.1, v0.75.0...
3 months ago

What's New

Closed-loop reward-hacking auto-mitigation — the GRPO/PPO trainer now detects reward hacking mid-run and self-corrects, instead of only halting. No OSS RLHF library (TRL / Unsloth / Axolotl / OpenRLHF / verl) closes this loop today.

Set training.reward_hack_mitigation (or soup train --reward-hack-mitigation) to one of four modes on a GRPO/PPO run (requires a reward_hack_detector):

  • log_only — instrument only: append a per-step mitigation_log.jsonl (drop_pct, verdict, reward mean/std, completion-length trend, repetition) and never touch training.
  • kl_control — a reversible bang-bang + hysteresis controller: when the hacking signal trips, raise the KL coefficient β (geometric, clamped to [floor, ceil], never crossing 0); relax it when the signal recovers. Dwell + release-patience prevent flapping; a multi-signal vote fuses the detector drop with a length-trend and a repetition signal.
  • pid_lagrangian — a PID-Lagrangian controller (Stooke et al.) that holds the hacking signal at a target, plus an escalation ladder: raise β → roll back to the last-good RL checkpoint → early-stop.
  • Anti-gaming hardening (Stage 3): per-signal EMA/median smoothing, conservative-on-disagreement voting, a reward-distribution-drift guard, and optional bounded reward shaping on the gamed proxy (length / repetition / sentinel). A plain-English give-up explanation is logged on early-stop.

Validated live on SmolLM2-135M + a synthetic length-hacking task (single RTX 3050): all four stages pass, including a real mid-run rollback to a last-good checkpoint.

Also: ready-made qwen2.5-coder-7b-sft recipe for Qwen/Qwen2.5-Coder-7B-Instruct (catalog 133 → 134, #285 by @Deadpool2000).

Install / Upgrade

pip install --upgrade soup-cli          # core CLI
pip install --upgrade 'soup-cli[train]' # + torch / transformers / trl / peft for training

Security

  • RLCheckpointCallback.restore_checkpoint / save_checkpoint refuse a symlinked optimizer.pt — torch.load(weights_only=False) on an attacker-placed symlink in a shared checkpoint dir was an RCE vector.
  • Bool-before-int/float guards on every new reward_hack_* numeric field; reward_hack_signals bounded (max_length=4).
  • The mitigation log writer is cwd-contained with symlink-reject-on-rotate and secret redaction (mirrors TraceLogWriter).

Known Limitations

  • Proof-of-mechanism ONLY, not a production claim. The controller is validated on SmolLM2-135M + a synthetic length-hacking reward on one RTX 3050. Whether it suppresses hacking without collapsing true reward on 7B+ with a real reward model is unproven — a help wanted issue tracks community reproduction at scale.
  • PPO ships BETA. The RLSignalBuffer + kl_coef mutation are wired and unit-tested, but the on-GPU proof is GRPO-only. PPO reward-capture only wraps callable reward fns (nn.Module reward models are skipped).
  • The InfoRM top/bottom-half detector split is itself gameable — a policy could shift the reward distribution to fool it. Stage 3's opt-in reward-distribution-drift guard is a partial mitigation, not full robustness.
  • Mid-run mitigation, not a new sampling scheme — this corrects an existing GRPO/PPO run; it is not on-policy distillation or a rollout change.
  • MitigationLogWriter rotation TOCTOU — the symlink-check-then-rename window is the same acknowledged limitation inherited from trace_logger.py (requires local-process write access; documented, not newly introduced).

Full changelog: https://github.com/MakazhanAlpamys/Soup/blob/main/CHANGELOG.md

Don't miss a new Soup release

NewReleases is sending notifications on new releases.