What's New
GRPO / Reasoning Training (Phase 4)
Train reasoning models with Group Relative Policy Optimization — the approach behind DeepSeek-R1.
base: meta-llama/Llama-3.1-8B-Instruct
task: grpo
training:
grpo_beta: 0.1
num_generations: 4
reward_fn: accuracy
lora:
r: 64
quantization: 4bitNew features:
task: grpo— generates multiple completions per prompt, scores with reward functions, optimizes using group-relative advantages- Built-in reward functions:
accuracy(checks final answer via####/\boxed{}) andformat(checks<think>...</think>reasoning blocks) - Custom reward functions — point
reward_fnto any.pyfile with areward_fn()callable soup init --template reasoning— ready-to-use GRPO config- Sweep support for
grpo_beta,num_generations,reward_fn
Stats: 371 tests (+42 new), 33 test files, lint clean
Install / Upgrade
pip install --upgrade soup-cliFull Changelog: v0.4.1...v0.4.2