github MakazhanAlpamys/Soup v0.4.2
v0.4.2 — GRPO Reasoning Training

latest releases: v0.75.2, v0.75.1, v0.75.0...
6 months ago

What's New

GRPO / Reasoning Training (Phase 4)

Train reasoning models with Group Relative Policy Optimization — the approach behind DeepSeek-R1.

base: meta-llama/Llama-3.1-8B-Instruct
task: grpo

training:
  grpo_beta: 0.1
  num_generations: 4
  reward_fn: accuracy
  lora:
    r: 64
  quantization: 4bit

New features:

  • task: grpo — generates multiple completions per prompt, scores with reward functions, optimizes using group-relative advantages
  • Built-in reward functions: accuracy (checks final answer via #### / \boxed{}) and format (checks <think>...</think> reasoning blocks)
  • Custom reward functions — point reward_fn to any .py file with a reward_fn() callable
  • soup init --template reasoning — ready-to-use GRPO config
  • Sweep support for grpo_beta, num_generations, reward_fn

Stats: 371 tests (+42 new), 33 test files, lint clean

Install / Upgrade

pip install --upgrade soup-cli

Full Changelog: v0.4.1...v0.4.2

Don't miss a new Soup release

NewReleases is sending notifications on new releases.