Tagline: Four v0.52.0 deferred stubs lifted into real trainer wrappers — knowledge distillation, sequence classification, EBFT / GDPO loss kernels, and gpt-oss-style reasoning_effort system-prompt injection.
What's New
soup trainwithtask: distill. NewDistillTrainerWrapper: student + frozen teacher both load viaAutoModelForCausalLM(separatetrust_remote_coderesolution per model), KL / forward_KL / reverse_KL / JS divergence kernels scaled bytemperature**2per the Hinton paper. Device-bridge survives HF Trainer's auto-CUDA promotion on a CPU-tagged run. UsesDataCollatorForSeq2Seq(label_pad_token_id=-100)so variable-length pre-tokenised loss-masked rows batch correctly.soup trainwithtask: classifier | reranker | cross_encoder. NewClassifierTrainerWrapper:AutoModelForSequenceClassificationwithnum_labels+label_names, auto-routes single-label / multi-label fromtcfg.classifier_kind. Multi-label string labels resolved via thelabel_namesmap with a 1024-entry cap + dedup. Training Setup Panel rendersHead: num_labels=N, kind=...instead of irrelevant LoRA r/alpha lines.- EBFT structured / strided + GDPO standard / length_normalized / margin loss kernels.
apply_ebft_loss/apply_gdpo_lossexit the v0.52.0NotImplementedErrorstubs with finite-only-input guards.attach_ebft_compute_loss(SFT) /attach_gdpo_compute_loss(DPO) wrapTrainer.compute_lossidempotently — re-attach is a no-op via a marker attribute. Auto-fires when the corresponding*_variantfield is set. - gpt-oss
reasoning_effort+train_on_eot.apply_reasoning_effort_prefix(messages, level)injects<|reasoning_effort|>{low,medium,high}<|/reasoning_effort|>into the system turn (creates one if absent), returning a new list.build_assistant_only_labels(train_on_eot=True)keeps the EOT/EOS token unmasked so the model learns when to stop. Both gated to the SFT-family at config-load. - +120 net new tests (7722 → 7842) in
test_v0532.py. Five review agents ran; every CRITICAL / HIGH / MEDIUM / LOW finding fixed or documented as a known limitation. Two real bugs surfaced and were fixed during the live CPU smoke (collator label padding + teacher / student device mismatch), both with source-level regression guards.
Install / Upgrade
pip install --upgrade soup-cliSecurity
- Separate
trust_remote_coderesolution for student vs teacher inDistillTrainerWrapper— a malicious teacher cannot piggy-back on the student's opt-in. - 1024-entry cap on multi-label rows in
ClassifierTrainerWrapper._normalise_labeldefeats OOM via crafted JSONL. - EBFT / GDPO kernels enforce finite-only tensor inputs (
torch.isfiniteguard) +math.isfiniteon scalar params — NaN / Inf would silently corrupt training otherwise. dpo_margindefaults toNone(not0.0) so the GDPOmarginvariant raises when the operator forgot to set the margin instead of silently producing a meaningless gradient.- Attach hooks (
attach_ebft_compute_loss,attach_gdpo_compute_loss) idempotent via marker attribute — re-attach is verified safe by a dedicated test class. - Distill teacher frozen via
requires_grad_(False)+.eval()immediately after load — never participates in gradient computation.
Known Limitations
- #71 TinyLlama-1.1B ONNX full export needs ≥16 GB free RAM.
optimum.main_exporttraces and serialises the model file successfully, but the post-processonnx.load(load_external_data=True)step needs to hold the full 4.4 GB fp32 weight blob in RAM alongside the serialised file. Tiny-gpt2 smoke confirms pipeline integrity. Host-RAM-bound, not a Soup bug. - Distillation supports same-tokenizer pairs only.
DistillTrainerWrapperassumes student and teacher share a tokenizer + vocab (logits aligned column-wise for KL). Cross-tokenizer distillation (e.g. Llama → Qwen) would need a projection or sequence-level loss — tracked for a follow-up. - Classifier wrapper has no LoRA path.
AutoModelForSequenceClassificationtrains the full head + base; LoRA-style classifier finetuning is a follow-up. - EBFT / GDPO auto-attach only fires when the corresponding
*_variantfield is set. Manualattach_ebft_compute_loss(trainer, tcfg)invocation from custom training loops is supported (idempotent), but plain SFT/DPO without the variant set never touches the hooks. reasoning_effortinjection happens at data-prep time. Changingtcfg.reasoning_effortbetween runs requires re-rendering the dataset (HF datasets caches the rendered chat template).
Full changelog: v0.53.1...v0.53.2