github MakazhanAlpamys/Soup v0.40.5
v0.40.5 — Quant Menu non-SFT

latest releases: v0.75.2, v0.75.1, v0.75.0...
4 months ago

What's New

Quant Menu non-SFT — closes the v0.38.0 known gap (#66). The seven train-time quantization formats (GPTQ / AWQ / HQQ:Nbit / AQLM / EETQ / MXFP4 / FP8) now work on every transformer-backend trainer, not just SFT.

  • Seven quant formats × 11 trainers — DPO / GRPO / KTO / ORPO / SimPO / IPO / PPO / RewardModel / Pretrain / Embedding / BCO can now load pre-quantized checkpoints (TheBloke/Llama-2-7B-GPTQ, etc.) and train LoRA adapters on top. The compatibility-matrix cross-validators (HQQ × ZeRO-3 reject, EETQ × FSDP reject, AQLM × FSDP reject) already gate at config-load and apply task-agnostic.
  • PPO reward model loads quantized too — _load_reward_model accepts an optional tcfg kwarg and routes the reward checkpoint through the same Quant Menu loader as the policy. A GPTQ-policy + GPTQ-reward PPO run no longer silently OOMs in fp16.
  • Defence-in-depth on reward_model field — TrainingConfig.reward_model field validator now rejects null bytes and caps length at 512 chars at config-load, matching the policy already applied to cfg.base.
  • +131 net new tests (4930 → 5061) across the schema-gate widening matrix (11 tasks × 7 formats), MLX rejection regression per task, source-level invariants asserting each non-SFT trainer no longer carries the legacy BitsAndBytesConfig(load_in_4bit=True ...) literal, and a live mock-based dispatch test for _load_reward_model.

Install / Upgrade

pip install --upgrade soup-cli

Security

  • TrainingConfig.reward_model field validator: null-byte rejection + 512-char cap at config-load (defence-in-depth ahead of the runtime null-byte check in quant_menu._check_local_marker).
  • All 11 non-SFT trainers now route quantized loads through build_quantization_config_for_loader, inheriting the v0.38.0 hardening (GPTQ quantize_config.json probe, AWQ quant_config.json probe, HQQ nbits allowlist, AQLM 2-bit lock, EETQ 8-bit lock, MXFP4 BNB validation, FP8 dequant).
  • prepare_model_for_kbit_training tuple widened to include mxfp4 so the BNB MXFP4 path correctly runs through kbit-prep instead of silently falling through.

Known Limitations

  1. Vision / audio modality + Quant Menu still rejected by the modality gate. The vision and audio paths in sft.py retain inline BitsAndBytesConfig blocks because they need vision-specific kwargs (LLaVA processor, etc.) that the unified loader does not yet thread. Multi-modal Quant Menu wiring is a follow-up item.
  2. Autopilot's quantization picker still recommends only 4bit / 8bit / none. Now that Quant Menu is wired across all transformer trainers, a future patch can teach Autopilot to recommend gptq / awq when base already points at a pre-quantized checkpoint.
  3. tcfg.reward_model is null-byte and length-validated at schema load but not path-containment-checked (is_under_cwd). This is consistent with how cfg.base is treated — both can be HF repo IDs or absolute local paths. The Quant Menu loader's read-only os.path.isfile probe is the only filesystem touch and uses os.path.basename in error messages so absolute paths are not leaked.

Don't miss a new Soup release

NewReleases is sending notifications on new releases.