github MakazhanAlpamys/Soup v0.71.30
v0.71.30 — PRM-guided GRPO + bundled rollout envs

latest releases: v0.75.2, v0.75.1, v0.75.0...
3 months ago

What's New

Process-supervised RL: PRM-guided GRPO. Use a trained Process Reward Model to score each reasoning step of a GRPO completion — the o1-era training signal — plus bundled toy environments so the openenv rollout path runs out-of-the-box.

  • PRM as the GRPO reward. Point training.prm_reward at a PRM you trained with soup train task=prm; it splits each completion into steps, scores every step with the PRM's reward head, and folds them (min / prod / last) into one reward that GRPO optimises — replacing reward_fn. It rides the reward-shaping + reward-hack-mitigation seam, so the v0.71.26 controller still watches it. Cross-validators gate task='grpo' + backend='transformers' + modality='text'.
  • Bundled rollout environments. soup_cli.envs.calculator / retrieval_qa / guess_number each expose a rollout(prompts) entry point; wire one with rollout_backend=openenv + rollout_func=soup_cli.envs.calculator:rollout. Three ready-made recipes ship it (grpo-env-calculator / grpo-env-retrieval-qa / grpo-env-guess-number); catalog 134 → 137.
task: grpo
training:
  prm_reward: ./my-prm            # a `soup train task=prm` checkpoint dir
  prm_aggregate: min              # weakest-link (default) | prod | last

Fixed

  • soup train task=prm producer conformance (surfaced by the live smoke): the PRM trainer now casts its reward head to the base-model dtype (bf16 CUDA runs previously crashed on the first compute_loss), saves the tokenizer alongside the model (so a PRM checkpoint is loadable standalone), and returns the standard trainer-result shape (previously the CLI crashed with a KeyError: 'initial_loss' right after saving).

Install / Upgrade

pip install -U soup-cli

Security

  • The PRM path is containment-checked (realpath + commonpath under cwd) before loading; reward-head weights load via safetensors.safe_open (no pickle); trust_remote_code defaults to False via the standard probe. The PRM-active console notice is rich.markup.escape'd against markup injection from a crafted config path. Bounded step count / per-step chars / total input tokens on the reward forward pass.

Known Limitations

  • Proof-of-mechanism only — validated on SmolLM2-135M with a tiny synthetic PRM and synthetic reward on CPU; not a production reward-model claim. Scale validation (7B+, real reward models) is help-wanted: #286.
  • Bundled environments are deterministic single-shot seeders, not interactive multi-turn model-in-the-loop episodes — the live openenv contract passes only the seed prompts, not the model.
  • Step split is a newline heuristic; the prompt context is rendered by joining message contents (matching the PRM's plaintext training field), not the tokenizer chat template the policy model sees.
  • PRM completions are scored one forward pass each (no batching) — fine for tiny models; batching is a future optimisation.
  • prm_aggregate='prod' assumes calibrated [0,1] step scores — Soup's PRM head is unconstrained-MSE-trained, so prod can blow up on uncalibrated labels; the default min is the safe choice.

Full history: CHANGELOG.md.

🤖 Generated with Claude Code

Don't miss a new Soup release

NewReleases is sending notifications on new releases.