github MakazhanAlpamys/Soup v0.53.2
v0.53.2 — Modality II live trainers

latest releases: v0.75.2, v0.75.1, v0.75.0...
4 months ago

Tagline: Four v0.52.0 deferred stubs lifted into real trainer wrappers — knowledge distillation, sequence classification, EBFT / GDPO loss kernels, and gpt-oss-style reasoning_effort system-prompt injection.

What's New

  • soup train with task: distill. New DistillTrainerWrapper: student + frozen teacher both load via AutoModelForCausalLM (separate trust_remote_code resolution per model), KL / forward_KL / reverse_KL / JS divergence kernels scaled by temperature**2 per the Hinton paper. Device-bridge survives HF Trainer's auto-CUDA promotion on a CPU-tagged run. Uses DataCollatorForSeq2Seq(label_pad_token_id=-100) so variable-length pre-tokenised loss-masked rows batch correctly.
  • soup train with task: classifier | reranker | cross_encoder. New ClassifierTrainerWrapper: AutoModelForSequenceClassification with num_labels + label_names, auto-routes single-label / multi-label from tcfg.classifier_kind. Multi-label string labels resolved via the label_names map with a 1024-entry cap + dedup. Training Setup Panel renders Head: num_labels=N, kind=... instead of irrelevant LoRA r/alpha lines.
  • EBFT structured / strided + GDPO standard / length_normalized / margin loss kernels. apply_ebft_loss / apply_gdpo_loss exit the v0.52.0 NotImplementedError stubs with finite-only-input guards. attach_ebft_compute_loss (SFT) / attach_gdpo_compute_loss (DPO) wrap Trainer.compute_loss idempotently — re-attach is a no-op via a marker attribute. Auto-fires when the corresponding *_variant field is set.
  • gpt-oss reasoning_effort + train_on_eot. apply_reasoning_effort_prefix(messages, level) injects <|reasoning_effort|>{low,medium,high}<|/reasoning_effort|> into the system turn (creates one if absent), returning a new list. build_assistant_only_labels(train_on_eot=True) keeps the EOT/EOS token unmasked so the model learns when to stop. Both gated to the SFT-family at config-load.
  • +120 net new tests (7722 → 7842) in test_v0532.py. Five review agents ran; every CRITICAL / HIGH / MEDIUM / LOW finding fixed or documented as a known limitation. Two real bugs surfaced and were fixed during the live CPU smoke (collator label padding + teacher / student device mismatch), both with source-level regression guards.

Install / Upgrade

pip install --upgrade soup-cli

Security

  • Separate trust_remote_code resolution for student vs teacher in DistillTrainerWrapper — a malicious teacher cannot piggy-back on the student's opt-in.
  • 1024-entry cap on multi-label rows in ClassifierTrainerWrapper._normalise_label defeats OOM via crafted JSONL.
  • EBFT / GDPO kernels enforce finite-only tensor inputs (torch.isfinite guard) + math.isfinite on scalar params — NaN / Inf would silently corrupt training otherwise.
  • dpo_margin defaults to None (not 0.0) so the GDPO margin variant raises when the operator forgot to set the margin instead of silently producing a meaningless gradient.
  • Attach hooks (attach_ebft_compute_loss, attach_gdpo_compute_loss) idempotent via marker attribute — re-attach is verified safe by a dedicated test class.
  • Distill teacher frozen via requires_grad_(False) + .eval() immediately after load — never participates in gradient computation.

Known Limitations

  • #71 TinyLlama-1.1B ONNX full export needs ≥16 GB free RAM. optimum.main_export traces and serialises the model file successfully, but the post-process onnx.load(load_external_data=True) step needs to hold the full 4.4 GB fp32 weight blob in RAM alongside the serialised file. Tiny-gpt2 smoke confirms pipeline integrity. Host-RAM-bound, not a Soup bug.
  • Distillation supports same-tokenizer pairs only. DistillTrainerWrapper assumes student and teacher share a tokenizer + vocab (logits aligned column-wise for KL). Cross-tokenizer distillation (e.g. Llama → Qwen) would need a projection or sequence-level loss — tracked for a follow-up.
  • Classifier wrapper has no LoRA path. AutoModelForSequenceClassification trains the full head + base; LoRA-style classifier finetuning is a follow-up.
  • EBFT / GDPO auto-attach only fires when the corresponding *_variant field is set. Manual attach_ebft_compute_loss(trainer, tcfg) invocation from custom training loops is supported (idempotent), but plain SFT/DPO without the variant set never touches the hooks.
  • reasoning_effort injection happens at data-prep time. Changing tcfg.reasoning_effort between runs requires re-rendering the dataset (HF datasets caches the rendered chat template).

Full changelog: v0.53.1...v0.53.2

Don't miss a new Soup release

NewReleases is sending notifications on new releases.