github MakazhanAlpamys/Soup v0.71.31
v0.71.31 — Judge-in-the-loop suite

latest releases: v0.75.2, v0.75.1, v0.75.0...
3 months ago

Put an LLM judge in the loop across the whole workflow: train against a judge, mine winners from a base model, grow instruction diversity, and decide SHIP with a true pairwise judge win-rate.

What's New

  • task='online_dpo' — Online DPO (wraps TRL OnlineDPOTrainer): the model generates two completions per prompt on-policy each step and a judge (a pairwise LLM judge over your local ollama / OpenAI-compatible endpoint) — or a reward_model — picks the winner. Set training.online_dpo_judge: "ollama://llama3.1" (or reward_model, exactly one), online_dpo_loss_type: sigmoid|ipo, online_dpo_max_new_tokens. Recipe: online-dpo-smollm2-135m. Adapts to the installed TRL — pairwise judge on trl 0.19.x, a pointwise reward function (the same JudgeEvaluator best-of-N uses) on trl 1.x, which removed pairwise judges.
  • soup data best-of-n — Best-of-N rejection sampling (BOND-lite): sample N completions from --base locally, a --judge scores each, and the winner becomes an SFT row. --emit-pairs also writes winner-vs-loser DPO pairs.
  • soup data evolve — WizardLM Evol-Instruct (depth / breadth) over an ollama / vllm provider — completing the synthetic-data suite (Magpie / Forge / Persona / evolve).
  • soup ship --task-mode pairwise (#284) — a true swap-debiased judge win-rate as the ship leg-1 task-win, fused with the catastrophic-forgetting guard into one SHIP / DON'T-SHIP verdict.

Install / Upgrade

pip install --upgrade soup-cli            # core
pip install --upgrade 'soup-cli[train]'   # + torch/transformers/peft/trl for training

Security

  • soup data best-of-n / evolve write outputs via atomic mkstemp + os.replace (re-validated cwd containment), closing the TOCTOU symlink-swap window between the containment check and the write.
  • All judge / provider URLs are SSRF-validated; model loads probe trust_remote_code; numeric inputs (n, rounds, tokens) are bounded.

Known Limitations

  • Proof-of-mechanism only — online-DPO validated with a synthetic length-preferring judge on SmolLM2-135M; not a production RLHF claim (scale help wanted: #286).
  • Per-version judge behaviour — on trl 0.19.x the online-DPO judge is a swap-debiased pairwise comparison; on trl 1.x (pairwise judges removed) the same judge model is used pointwise. Documented behaviour difference.
  • Pairwise judge = 2× calls per pair — soup ship --task-mode pairwise and the 0.19.x online-DPO judge judge both orders for swap-debias.
  • best-of-N sampling loads the full model locally (bounded by your box; ≤3B validated); a provider-based sampler is future work.

Don't miss a new Soup release

NewReleases is sending notifications on new releases.