github MakazhanAlpamys/Soup v0.53.4
v0.53.4 — Long Context + Architecture

latest releases: v0.75.2, v0.75.1, v0.75.0...
4 months ago

What's New

Six closes lifting the v0.49.0 LongLoRA hardening + v0.41.0 LLaMA Pro deferred stubs, plus a UX upgrade to the CUDA-OOM friendly message.

  • LLaMA Pro is live. training.expand_layers: N now actually appends N zero-initialised decoder blocks (deep-copied from the tail, residual projections zeroed so the appended block acts as identity per the LLaMA Pro paper §3.1). Pair with freeze_trainable_layers: N to train only the new blocks. Wired into SFT and Pretrain trainers via the centralised apply_block_expansion_if_configured helper (mirrors v0.40.6 peft_wiring centralisation policy).
  • LongLoRA arch allowlist expanded. use_longlora: true now accepts Llama / CodeLlama / Mistral / Qwen / Phi base models via new word-boundary helpers (is_mistral_model, is_qwen_model, is_phi_model). Mixtral is intentionally excluded — its MoE attention requires a dedicated helper still tracked for a future release.
  • LongLoRA + FlashAttention v3 schema reject. New flash_attn.is_flash_attn_v3_available() defensive probe — when FA-v3 is installed the schema now rejects use_longlora: true at config load with an actionable error. S² shifted-sparse + FA-v3 custom-mask both rewrite the attention kernel; allowing both would silently corrupt outputs.
  • Llama 3.1 RoPE auto-detect. Omit rope_scaling_type from your YAML on a Llama 3.1 base and apply_long_context_config now reads model.config.rope_scaling. If it carries a llama3 block, the long-context path picks LLAMA3_DEFAULT_* instead of falling through to dynamic. Explicit caller picks still win.
  • CUDA-OOM hint upgrade. format_friendly_error now points users at the exact CLI flags — --batch-size <half> and --grad-accum <double> to preserve effective batch size — before the legacy quantization: 4bit fallback.

Install / Upgrade

pip install --upgrade soup-cli

Or for the latest dev:

pip install git+https://github.com/MakazhanAlpamys/Soup.git@v0.53.4

Security

Six closes maintain the project's hardening invariants — see SECURITY.md for the full per-fix breakdown. Highlights from the v0.53.4 review pipeline:

  • Defensive input surface on every new helper. _check_model_name rejects bool BEFORE the isinstance(str) check (bool is a subclass of int and would otherwise fall through silently). is_supported_longlora_arch returns False on non-string input rather than propagating TypeError.
  • 64-char base echo truncation in LongLoRA error messages via new _truncate_for_message helper — defends against adversarial / long bases bloating stderr + log files (mirrors the v0.53.3 validate_vision_grpo_compat redaction).
  • task / backend null-byte rejection in validate_longlora_compat — defends against null bytes in user-controlled YAML leaking literally into error messages and downstream log files.
  • Explicit is None in _get_layers_module — defends against nn.Module.__bool__ overrides on subclasses that would otherwise silently fall through to the wrong code path.
  • warnings.warn on non-Llama-shaped expansion — _zero_init_block_residual returns bool + the caller emits a runtime warning when neither standard projection matches; non-Llama-shaped architectures still train but lose the LLaMA Pro identity-init guarantee.

Test Surface

  • 7879 → 7935 tests (+56 net; +49 in new tests/test_v0534.py).
  • Lint clean; full four-agent review pipeline ran (python / code / security / tdd).
  • Real-model CPU smoke verified on transformers.LlamaForCausalLM: 4 → 6 layers, down_proj + o_proj actually zeroed on PyTorch tensors, old blocks frozen + new blocks trainable, forward pass produces finite logits.

Known Limitations

  • LongLoRA S² forward override remains deferred. The schema gate is hardened; live LlamaAttention.forward monkeypatch lands in a follow-up release.
  • Mixtral excluded from LongLoRA allowlist. MoE attention forward signature differs from standard Mistral; a dedicated helper is tracked separately.
  • Block-expansion zero-init covers Llama-shaped blocks only. Non-standard architectures (e.g. Falcon's dense_4h_to_h) emit a warnings.warn and continue — the appended block is still trainable but lacks the identity-init guarantee.
  • Llama 3.1 RoPE auto-detect only fires when caller explicitly passes rope_scaling_type=None. Explicit picks always win; this is intended for back-compat.
  • #74 live HF push QA was deferred to a contributor with private HF credentials — pipeline integrity verified via --help + plumbing only.

Full changelog

v0.53.3...v0.53.4

Don't miss a new Soup release

NewReleases is sending notifications on new releases.