What's New
Six closes lifting the v0.49.0 LongLoRA hardening + v0.41.0 LLaMA Pro deferred stubs, plus a UX upgrade to the CUDA-OOM friendly message.
- LLaMA Pro is live.
training.expand_layers: Nnow actually appends N zero-initialised decoder blocks (deep-copied from the tail, residual projections zeroed so the appended block acts as identity per the LLaMA Pro paper §3.1). Pair withfreeze_trainable_layers: Nto train only the new blocks. Wired into SFT and Pretrain trainers via the centralisedapply_block_expansion_if_configuredhelper (mirrors v0.40.6peft_wiringcentralisation policy). - LongLoRA arch allowlist expanded.
use_longlora: truenow accepts Llama / CodeLlama / Mistral / Qwen / Phi base models via new word-boundary helpers (is_mistral_model,is_qwen_model,is_phi_model). Mixtral is intentionally excluded — its MoE attention requires a dedicated helper still tracked for a future release. - LongLoRA + FlashAttention v3 schema reject. New
flash_attn.is_flash_attn_v3_available()defensive probe — when FA-v3 is installed the schema now rejectsuse_longlora: trueat config load with an actionable error. S² shifted-sparse + FA-v3 custom-mask both rewrite the attention kernel; allowing both would silently corrupt outputs. - Llama 3.1 RoPE auto-detect. Omit
rope_scaling_typefrom your YAML on a Llama 3.1 base andapply_long_context_confignow readsmodel.config.rope_scaling. If it carries allama3block, the long-context path picksLLAMA3_DEFAULT_*instead of falling through todynamic. Explicit caller picks still win. - CUDA-OOM hint upgrade.
format_friendly_errornow points users at the exact CLI flags —--batch-size <half>and--grad-accum <double>to preserve effective batch size — before the legacyquantization: 4bitfallback.
Install / Upgrade
pip install --upgrade soup-cliOr for the latest dev:
pip install git+https://github.com/MakazhanAlpamys/Soup.git@v0.53.4Security
Six closes maintain the project's hardening invariants — see SECURITY.md for the full per-fix breakdown. Highlights from the v0.53.4 review pipeline:
- Defensive input surface on every new helper.
_check_model_namerejectsboolBEFORE theisinstance(str)check (bool is a subclass of int and would otherwise fall through silently).is_supported_longlora_archreturns False on non-string input rather than propagating TypeError. - 64-char
baseecho truncation in LongLoRA error messages via new_truncate_for_messagehelper — defends against adversarial / long bases bloating stderr + log files (mirrors the v0.53.3validate_vision_grpo_compatredaction). task/backendnull-byte rejection invalidate_longlora_compat— defends against null bytes in user-controlled YAML leaking literally into error messages and downstream log files.- Explicit
is Nonein_get_layers_module— defends againstnn.Module.__bool__overrides on subclasses that would otherwise silently fall through to the wrong code path. warnings.warnon non-Llama-shaped expansion —_zero_init_block_residualreturns bool + the caller emits a runtime warning when neither standard projection matches; non-Llama-shaped architectures still train but lose the LLaMA Pro identity-init guarantee.
Test Surface
- 7879 → 7935 tests (+56 net; +49 in new
tests/test_v0534.py). - Lint clean; full four-agent review pipeline ran (python / code / security / tdd).
- Real-model CPU smoke verified on
transformers.LlamaForCausalLM: 4 → 6 layers,down_proj+o_projactually zeroed on PyTorch tensors, old blocks frozen + new blocks trainable, forward pass produces finite logits.
Known Limitations
- LongLoRA S² forward override remains deferred. The schema gate is hardened; live
LlamaAttention.forwardmonkeypatch lands in a follow-up release. - Mixtral excluded from LongLoRA allowlist. MoE attention forward signature differs from standard Mistral; a dedicated helper is tracked separately.
- Block-expansion zero-init covers Llama-shaped blocks only. Non-standard architectures (e.g. Falcon's
dense_4h_to_h) emit awarnings.warnand continue — the appended block is still trainable but lacks the identity-init guarantee. - Llama 3.1 RoPE auto-detect only fires when caller explicitly passes
rope_scaling_type=None. Explicit picks always win; this is intended for back-compat. - #74 live HF push QA was deferred to a contributor with private HF credentials — pipeline integrity verified via
--help+ plumbing only.