github MakazhanAlpamys/Soup v0.9.0
v0.9.0 — PPO / Full RLHF Pipeline

latest releases: v0.75.2, v0.75.1, v0.75.0...
6 months ago

What's New

PPO / Full RLHF Pipeline (Phase 10)

Three-stage RLHF training: SFT → Reward Model → PPO. Completes the full alignment pipeline.

New Features

  • task: ppo — Proximal Policy Optimization with TRL's PPOTrainer
    • Manual training loop: generate completions → score with reward → PPO step
    • Two reward sources: pre-trained reward model (reward_model) and/or callable reward function (reward_fn)
    • Config: ppo_epochs, ppo_clip_ratio, ppo_kl_penalty
  • task: reward_model — Train reward models from preference data
    • Uses TRL's RewardTrainer with AutoModelForSequenceClassification
    • Input format: same as DPO (prompt/chosen/rejected)
  • soup init --template rlhf — RLHF template with PPO config
  • Sweep support — ppo_epochs, ppo_clip_ratio, ppo_kl_penalty, reward_model shortcuts
  • Compatible with Unsloth backend, QAT, DeepSpeed

Tests

  • 51 new tests (611 total across 40 test files)
  • All tests passing, ruff clean

Install / Upgrade

pip install --upgrade soup-cli

Full Changelog

v0.8.0...v0.9.0

Don't miss a new Soup release

NewReleases is sending notifications on new releases.