What's New
Performance Optimizations
- Liger Kernel — fused CUDA ops (RMSNorm, SwiGLU, CrossEntropy, RoPE) for 20-60% memory savings.
use_liger: true - FlashAttention v2/v3 — auto-detects best attention implementation.
use_flash_attn: true - FSDP2 — PyTorch-native distributed training.
--fsdp fsdp_full_shard|fsdp_shard_grad|fsdp_full_offload - Gradient checkpointing —
gradient_checkpointing: truefor memory-efficient long sequences
Long-Context Fine-Tuning (128k+)
- RoPE scaling —
rope_scaling_type: dynamic(also: linear, yarn, longrope) - Ring FlashAttention — sequence parallelism across GPUs.
use_ring_attention: true - New template —
soup init --template longcontext
Security
- rope_scaling_type Literal constraint, max_length bounds (64-1M), FSDP key allowlist, Liger exception narrowing
Testing
- 91 new tests (1182 total), 49 test files, 58.5% coverage
Install / Upgrade
pip install -U soup-cli
pip install 'soup-cli[liger]' # Liger Kernel
pip install 'soup-cli[ring-attn]' # Ring FlashAttention
pip install flash-attn --no-build-isolation # FlashAttentionFull Changelog: v0.14.3...v0.15.0