What's New
PPO / Full RLHF Pipeline (Phase 10)
Three-stage RLHF training: SFT → Reward Model → PPO. Completes the full alignment pipeline.
New Features
task: ppo— Proximal Policy Optimization with TRL's PPOTrainer- Manual training loop: generate completions → score with reward → PPO step
- Two reward sources: pre-trained reward model (
reward_model) and/or callable reward function (reward_fn) - Config:
ppo_epochs,ppo_clip_ratio,ppo_kl_penalty
task: reward_model— Train reward models from preference data- Uses TRL's RewardTrainer with
AutoModelForSequenceClassification - Input format: same as DPO (prompt/chosen/rejected)
- Uses TRL's RewardTrainer with
soup init --template rlhf— RLHF template with PPO config- Sweep support —
ppo_epochs,ppo_clip_ratio,ppo_kl_penalty,reward_modelshortcuts - Compatible with Unsloth backend, QAT, DeepSpeed
Tests
- 51 new tests (611 total across 40 test files)
- All tests passing, ruff clean
Install / Upgrade
pip install --upgrade soup-cli