What's New
Two surgical fixes from the v0.50.0 GRPO Plus deferred-stub family. The four larger items (stability callback, variant losses, PRM trainer, multi-objective preference live combine) are scope-deferred to v0.53.4 — each warrants a focused release per the v0.40.x stub-then-live cadence.
grpo_fp16: truenow wires through GRPOTrainerWrapper. New_build_precision_kwargs(self) -> dict[str, bool]returns the right{fp16, bf16}per(device, grpo_fp16)matrix: non-CUDA (CPU/MPS/XPU) → both False (HF Trainer's mixed-precision kwargs are CUDA-specific), CUDA +grpo_fp16=True→fp16=True, bf16=False(unsloth parity), default CUDA → bf16 (legacy v0.50.0 path).- Mutex cross-validator.
SoupConfig._validate_grpo_fp16_amp_exclusiverejects the silent footgungrpo_fp16=True+auto_mixed_precision=True. The validator short-circuits whentask != 'grpo'so the v0.50.0 stability task-gate diagnosis fires first, keeping the most actionable error in front of the user regardless of validator execution order. - Vision-GRPO base-model probe. New
KNOWN_VLM_REGEXcovers 10 VLM families (Qwen2-VL / Qwen2.5-VL / QVQ / Pixtral / InternVL / Llama-3.2-Vision / LLaVA / MiniCPM-V / Idefics / ShareGPT4V / Fuyu) with word-boundary anchors that reject substring noise like"my-pixtralish".is_known_vlm_basereturns False (never raises) on bad input.validate_vision_grpo_compataccepts a new optionalbasekwarg threaded fromSoupConfig._validate_vision_grpo. - Friendly schema rejection instead of a cryptic runtime
AttributeError: 'vision_tower': a YAML pairingvision_grpo: truewith a non-VLM checkpoint now fails at config-load with a message naming the expected families. Echoedbasevalue is truncated to 64 chars before serialisation, mirroring the v0.34.0crash.pyredaction policy. - +37 net new tests (7842 → 7879) in
tests/test_v0533.py. Four review agents (python / code / security / tdd) ran; every HIGH / MEDIUM / LOW finding fixed.
Install / Upgrade
pip install --upgrade soup-cliSecurity
- Vision-GRPO base probe rejects unknown VLMs at schema load (defense-in-depth — keeps the legacy attribute-error footgun off the runtime path).
- 64-char truncation on adversarial / oversized base values in the rejection error message — prevents log bloat and unredacted user input from surfacing in operator-facing tracebacks.
is_known_vlm_baseis defensive — returns False (never raises) on bool / non-string / empty / null-byte />_MAX_BASE_NAME_LEN=512, mirroring v0.39.0 / v0.44.0 / v0.49.0 model-detection policy.
Known Limitations
- Scope-deferred —
#127GRPOStabilityCallback (7 stability knobs),#1236 GRPO variant loss kernels (gspo / dapo / dr_grpo / bnpo / two_sided / rft),#126PRMTrainerWrapper, and#68PreferenceTrainerWrapper true weighted-loss combine all move to v0.53.4. Each requires deep TRL trainer subclassing and warrants its own focused release. - VLM allowlist is static name-regex — a legitimate VLM published under an org whose checkpoint name lacks any of the 10 tokens is rejected at schema-load; users must omit
vision_grpo: trueuntil a future release adds a runtimemodel.config.vision_configprobe. _build_precision_kwargsis GRPO-only — other RL trainers (PPO / RewardModel) follow their existing mixed-precision conventions; thegrpo_fp16flag is GRPO-gated by the v0.50.0 stability task-gate.