What's New
Process-supervised RL: PRM-guided GRPO. Use a trained Process Reward Model to score each reasoning step of a GRPO completion — the o1-era training signal — plus bundled toy environments so the openenv rollout path runs out-of-the-box.
- PRM as the GRPO reward. Point
training.prm_rewardat a PRM you trained withsoup train task=prm; it splits each completion into steps, scores every step with the PRM's reward head, and folds them (min/prod/last) into one reward that GRPO optimises — replacingreward_fn. It rides the reward-shaping + reward-hack-mitigation seam, so the v0.71.26 controller still watches it. Cross-validators gatetask='grpo'+backend='transformers'+modality='text'. - Bundled rollout environments.
soup_cli.envs.calculator/retrieval_qa/guess_numbereach expose arollout(prompts)entry point; wire one withrollout_backend=openenv+rollout_func=soup_cli.envs.calculator:rollout. Three ready-made recipes ship it (grpo-env-calculator/grpo-env-retrieval-qa/grpo-env-guess-number); catalog 134 → 137.
task: grpo
training:
prm_reward: ./my-prm # a `soup train task=prm` checkpoint dir
prm_aggregate: min # weakest-link (default) | prod | lastFixed
soup train task=prmproducer conformance (surfaced by the live smoke): the PRM trainer now casts its reward head to the base-model dtype (bf16 CUDA runs previously crashed on the firstcompute_loss), saves the tokenizer alongside the model (so a PRM checkpoint is loadable standalone), and returns the standard trainer-result shape (previously the CLI crashed with aKeyError: 'initial_loss'right after saving).
Install / Upgrade
pip install -U soup-cliSecurity
- The PRM path is containment-checked (
realpath+commonpathunder cwd) before loading; reward-head weights load viasafetensors.safe_open(no pickle);trust_remote_codedefaults toFalsevia the standard probe. The PRM-active console notice isrich.markup.escape'd against markup injection from a crafted config path. Bounded step count / per-step chars / total input tokens on the reward forward pass.
Known Limitations
- Proof-of-mechanism only — validated on SmolLM2-135M with a tiny synthetic PRM and synthetic reward on CPU; not a production reward-model claim. Scale validation (7B+, real reward models) is help-wanted: #286.
- Bundled environments are deterministic single-shot seeders, not interactive multi-turn model-in-the-loop episodes — the live openenv contract passes only the seed prompts, not the model.
- Step split is a newline heuristic; the prompt context is rendered by joining message contents (matching the PRM's plaintext training field), not the tokenizer chat template the policy model sees.
- PRM completions are scored one forward pass each (no batching) — fine for tiny models; batching is a future optimisation.
prm_aggregate='prod'assumes calibrated[0,1]step scores — Soup's PRM head is unconstrained-MSE-trained, soprodcan blow up on uncalibrated labels; the defaultminis the safe choice.
Full history: CHANGELOG.md.
🤖 Generated with Claude Code