What's New
Six new training-speed & memory features, task: sft wired in this release (multi-trainer support in v0.28.1):
- Cut Cross-Entropy (CCE) —
training.use_cut_ce: truefuses the LM-head + cross-entropy. Saves 8-24 GB VRAM on 128k-vocab models (Llama 3.1, Qwen2). Install:pip install 'soup-cli[cce]'. - FP8 training —
training.quantization_aware: "fp8"enables float8 matmuls on Hopper+ (H100/H200/B100/B200) viatorchao.float8. Booltruekeeps the legacy int8 QAT path. - Gradient checkpointing tiers —
training.gradient_checkpointing: selective | medium | full | auto.autopicks based on detected VRAM (<24 GB → full, 24-80 GB → medium, >80 GB → selective). - Kernel auto-composition —
training.kernel_auto_compose: trueenumerates baseline / Liger / FlashAttn / CCE combos and picks the fastest. - Cross-document attention masking —
training.packing_cross_doc_attn_mask: truewithpacking: trueblocks attention from crossing document boundaries in packed sequences. - Activation offloading —
training.activation_offloading: cpu | diskoffloads saved activations to RAM or a scratch file during the backward pass.
Install / Upgrade
pip install --upgrade soup-cli
# Optional extras for v0.28.0 features:
pip install 'soup-cli[cce]' # Cut Cross-Entropy (large-vocab save)
pip install 'soup-cli[qat]' # FP8 (torchao) — required for quantization_aware: "fp8"Security
quantization_aware: Union[bool, Literal[\"fp8\"]]— Pydantic rejects arbitrary strings (onlytrue/false/"fp8"). FP8 path requires CUDA + Hopper+ SM capability + transformers backend.gradient_checkpointing: Union[bool, Literal[\"selective\",\"medium\",\"full\",\"auto\"]]— rejects unknown tier strings; returns only HF-supported keys toTrainingArguments(no private markers leak).activation_offloadingLiteralcpu|disk— scratchsave_dircontainment-enforced viautils/paths.is_under_cwdbefore disk writes.torch.load(weights_only=True)prevents arbitrary Python deserialization on reload. TOCTOU closed betweenmkstempandtorch.saveby holding the fd open. Best-effort cleanup on context exit (handles SIGKILL mid-backward).kernel_picker.pick_best_kernelraisesValueErrorwhen all candidates lack a finitetime_ms— prevents silent promotion of an untimed combo.- Cut CE architecture detector matches on last path component only, so
deepseek-ai/...-phi-...org-prefix does not trigger a Phi patch on a DeepSeek model. build_cross_doc_masknumpy-vectorised (np.tril) to avoid O(seq_length²) pure-Python fill at the 1Mmax_lengthbound.@model_validatorgates:packing_cross_doc_attn_maskrequirespacking=true;use_cut_ce/quantization_aware="fp8"/kernel_auto_compose/activation_offloadingrequiretask=sft.
Known Limitations
- SFT-only wiring:
use_cut_ce,quantization_aware="fp8",kernel_auto_compose,activation_offloadingimplemented only inSFTTrainerWrapper. Config-load validator fails fast on non-SFT tasks. Multi-trainer wiring (DPO/GRPO/KTO/ORPO/SimPO/IPO/PPO/Pretrain/RewardModel/Embedding) tracked for v0.28.1. - Gradient checkpointing tier granularity:
selective/medium/autocurrently behave likefullat the HFTrainingArgumentslayer. Memory savings are real; the per-layer selectivity (attention-only hooks) ships in v0.28.1. - Kernel picker benchmark loop:
enumerate_kernel_combos+pick_best_kernelare pure helpers. The 2-3 step warm-up benchmark inside the trainer is deferred to v0.28.1 — v0.28.0 ships the picker API + schema flag only. - FP8 recipe: uses torchao's default tensorwise scaling. Row-wise + delayed-scaling recipes deferred to v0.28.1.
- Cross-doc attention mask wiring: uses TRL's
packing_strategy="attention_free"hint. Full collator integration (explicit block-diagonal mask passed to the attention kernel) tracked for v0.28.1. - License migration (MIT → Apache-2.0) applied 2026-04-21 — downstream redistributors must retain the
NOTICEfile per Apache-2.0 §4(d).
Tests: 2585 → 2685 (+100). CI: green across ubuntu / windows / macos × Python 3.9 / 3.11 / 3.12.