github MakazhanAlpamys/Soup v0.28.0
v0.28.0 — Training Speed & Memory

latest releases: v0.75.2, v0.75.1, v0.75.0...
5 months ago

What's New

Six new training-speed & memory features, task: sft wired in this release (multi-trainer support in v0.28.1):

  • Cut Cross-Entropy (CCE) — training.use_cut_ce: true fuses the LM-head + cross-entropy. Saves 8-24 GB VRAM on 128k-vocab models (Llama 3.1, Qwen2). Install: pip install 'soup-cli[cce]'.
  • FP8 training — training.quantization_aware: "fp8" enables float8 matmuls on Hopper+ (H100/H200/B100/B200) via torchao.float8. Bool true keeps the legacy int8 QAT path.
  • Gradient checkpointing tiers — training.gradient_checkpointing: selective | medium | full | auto. auto picks based on detected VRAM (<24 GB → full, 24-80 GB → medium, >80 GB → selective).
  • Kernel auto-composition — training.kernel_auto_compose: true enumerates baseline / Liger / FlashAttn / CCE combos and picks the fastest.
  • Cross-document attention masking — training.packing_cross_doc_attn_mask: true with packing: true blocks attention from crossing document boundaries in packed sequences.
  • Activation offloading — training.activation_offloading: cpu | disk offloads saved activations to RAM or a scratch file during the backward pass.

Install / Upgrade

pip install --upgrade soup-cli
# Optional extras for v0.28.0 features:
pip install 'soup-cli[cce]'   # Cut Cross-Entropy (large-vocab save)
pip install 'soup-cli[qat]'   # FP8 (torchao) — required for quantization_aware: "fp8"

Security

  • quantization_aware: Union[bool, Literal[\"fp8\"]] — Pydantic rejects arbitrary strings (only true / false / "fp8"). FP8 path requires CUDA + Hopper+ SM capability + transformers backend.
  • gradient_checkpointing: Union[bool, Literal[\"selective\",\"medium\",\"full\",\"auto\"]] — rejects unknown tier strings; returns only HF-supported keys to TrainingArguments (no private markers leak).
  • activation_offloading Literal cpu|disk — scratch save_dir containment-enforced via utils/paths.is_under_cwd before disk writes. torch.load(weights_only=True) prevents arbitrary Python deserialization on reload. TOCTOU closed between mkstemp and torch.save by holding the fd open. Best-effort cleanup on context exit (handles SIGKILL mid-backward).
  • kernel_picker.pick_best_kernel raises ValueError when all candidates lack a finite time_ms — prevents silent promotion of an untimed combo.
  • Cut CE architecture detector matches on last path component only, so deepseek-ai/...-phi-... org-prefix does not trigger a Phi patch on a DeepSeek model.
  • build_cross_doc_mask numpy-vectorised (np.tril) to avoid O(seq_length²) pure-Python fill at the 1M max_length bound.
  • @model_validator gates: packing_cross_doc_attn_mask requires packing=true; use_cut_ce / quantization_aware="fp8" / kernel_auto_compose / activation_offloading require task=sft.

Known Limitations

  • SFT-only wiring: use_cut_ce, quantization_aware="fp8", kernel_auto_compose, activation_offloading implemented only in SFTTrainerWrapper. Config-load validator fails fast on non-SFT tasks. Multi-trainer wiring (DPO/GRPO/KTO/ORPO/SimPO/IPO/PPO/Pretrain/RewardModel/Embedding) tracked for v0.28.1.
  • Gradient checkpointing tier granularity: selective / medium / auto currently behave like full at the HF TrainingArguments layer. Memory savings are real; the per-layer selectivity (attention-only hooks) ships in v0.28.1.
  • Kernel picker benchmark loop: enumerate_kernel_combos + pick_best_kernel are pure helpers. The 2-3 step warm-up benchmark inside the trainer is deferred to v0.28.1 — v0.28.0 ships the picker API + schema flag only.
  • FP8 recipe: uses torchao's default tensorwise scaling. Row-wise + delayed-scaling recipes deferred to v0.28.1.
  • Cross-doc attention mask wiring: uses TRL's packing_strategy="attention_free" hint. Full collator integration (explicit block-diagonal mask passed to the attention kernel) tracked for v0.28.1.
  • License migration (MIT → Apache-2.0) applied 2026-04-21 — downstream redistributors must retain the NOTICE file per Apache-2.0 §4(d).

Tests: 2585 → 2685 (+100). CI: green across ubuntu / windows / macos × Python 3.9 / 3.11 / 3.12.

Don't miss a new Soup release

NewReleases is sending notifications on new releases.