What's New
Closed-loop reward-hacking auto-mitigation — the GRPO/PPO trainer now detects reward hacking mid-run and self-corrects, instead of only halting. No OSS RLHF library (TRL / Unsloth / Axolotl / OpenRLHF / verl) closes this loop today.
Set training.reward_hack_mitigation (or soup train --reward-hack-mitigation) to one of four modes on a GRPO/PPO run (requires a reward_hack_detector):
log_only— instrument only: append a per-stepmitigation_log.jsonl(drop_pct, verdict, reward mean/std, completion-length trend, repetition) and never touch training.kl_control— a reversible bang-bang + hysteresis controller: when the hacking signal trips, raise the KL coefficient β (geometric, clamped to[floor, ceil], never crossing 0); relax it when the signal recovers. Dwell + release-patience prevent flapping; a multi-signal vote fuses the detector drop with a length-trend and a repetition signal.pid_lagrangian— a PID-Lagrangian controller (Stooke et al.) that holds the hacking signal at a target, plus an escalation ladder: raise β → roll back to the last-good RL checkpoint → early-stop.- Anti-gaming hardening (Stage 3): per-signal EMA/median smoothing, conservative-on-disagreement voting, a reward-distribution-drift guard, and optional bounded reward shaping on the gamed proxy (length / repetition / sentinel). A plain-English give-up explanation is logged on early-stop.
Validated live on SmolLM2-135M + a synthetic length-hacking task (single RTX 3050): all four stages pass, including a real mid-run rollback to a last-good checkpoint.
Also: ready-made qwen2.5-coder-7b-sft recipe for Qwen/Qwen2.5-Coder-7B-Instruct (catalog 133 → 134, #285 by @Deadpool2000).
Install / Upgrade
pip install --upgrade soup-cli # core CLI
pip install --upgrade 'soup-cli[train]' # + torch / transformers / trl / peft for trainingSecurity
RLCheckpointCallback.restore_checkpoint/save_checkpointrefuse a symlinkedoptimizer.pt—torch.load(weights_only=False)on an attacker-placed symlink in a shared checkpoint dir was an RCE vector.- Bool-before-int/float guards on every new
reward_hack_*numeric field;reward_hack_signalsbounded (max_length=4). - The mitigation log writer is cwd-contained with symlink-reject-on-rotate and secret redaction (mirrors
TraceLogWriter).
Known Limitations
- Proof-of-mechanism ONLY, not a production claim. The controller is validated on SmolLM2-135M + a synthetic length-hacking reward on one RTX 3050. Whether it suppresses hacking without collapsing true reward on 7B+ with a real reward model is unproven — a
help wantedissue tracks community reproduction at scale. - PPO ships BETA. The RLSignalBuffer +
kl_coefmutation are wired and unit-tested, but the on-GPU proof is GRPO-only. PPO reward-capture only wraps callable reward fns (nn.Module reward models are skipped). - The InfoRM top/bottom-half detector split is itself gameable — a policy could shift the reward distribution to fool it. Stage 3's opt-in reward-distribution-drift guard is a partial mitigation, not full robustness.
- Mid-run mitigation, not a new sampling scheme — this corrects an existing GRPO/PPO run; it is not on-policy distillation or a rollout change.
- MitigationLogWriter rotation TOCTOU — the symlink-check-then-rename window is the same acknowledged limitation inherited from
trace_logger.py(requires local-process write access; documented, not newly introduced).
Full changelog: https://github.com/MakazhanAlpamys/Soup/blob/main/CHANGELOG.md