Put an LLM judge in the loop across the whole workflow: train against a judge, mine winners from a base model, grow instruction diversity, and decide SHIP with a true pairwise judge win-rate.
What's New
task='online_dpo'— Online DPO (wraps TRLOnlineDPOTrainer): the model generates two completions per prompt on-policy each step and a judge (a pairwise LLM judge over your local ollama / OpenAI-compatible endpoint) — or areward_model— picks the winner. Settraining.online_dpo_judge: "ollama://llama3.1"(orreward_model, exactly one),online_dpo_loss_type: sigmoid|ipo,online_dpo_max_new_tokens. Recipe:online-dpo-smollm2-135m. Adapts to the installed TRL — pairwise judge on trl 0.19.x, a pointwise reward function (the sameJudgeEvaluatorbest-of-N uses) on trl 1.x, which removed pairwise judges.soup data best-of-n— Best-of-N rejection sampling (BOND-lite): sample N completions from--baselocally, a--judgescores each, and the winner becomes an SFT row.--emit-pairsalso writes winner-vs-loser DPO pairs.soup data evolve— WizardLM Evol-Instruct (depth / breadth) over an ollama / vllm provider — completing the synthetic-data suite (Magpie / Forge / Persona / evolve).soup ship --task-mode pairwise(#284) — a true swap-debiased judge win-rate as the ship leg-1 task-win, fused with the catastrophic-forgetting guard into one SHIP / DON'T-SHIP verdict.
Install / Upgrade
pip install --upgrade soup-cli # core
pip install --upgrade 'soup-cli[train]' # + torch/transformers/peft/trl for training
Security
soup data best-of-n/evolvewrite outputs via atomicmkstemp+os.replace(re-validated cwd containment), closing the TOCTOU symlink-swap window between the containment check and the write.- All judge / provider URLs are SSRF-validated; model loads probe
trust_remote_code; numeric inputs (n, rounds, tokens) are bounded.
Known Limitations
- Proof-of-mechanism only — online-DPO validated with a synthetic length-preferring judge on SmolLM2-135M; not a production RLHF claim (scale help wanted: #286).
- Per-version judge behaviour — on trl 0.19.x the online-DPO judge is a swap-debiased pairwise comparison; on trl 1.x (pairwise judges removed) the same judge model is used pointwise. Documented behaviour difference.
- Pairwise judge = 2× calls per pair —
soup ship --task-mode pairwiseand the 0.19.x online-DPO judge judge both orders for swap-debias. - best-of-N sampling loads the full model locally (bounded by your box; ≤3B validated); a provider-based sampler is future work.