What's New
Quant Menu non-SFT — closes the v0.38.0 known gap (#66). The seven train-time quantization formats (GPTQ / AWQ / HQQ:Nbit / AQLM / EETQ / MXFP4 / FP8) now work on every transformer-backend trainer, not just SFT.
- Seven quant formats × 11 trainers — DPO / GRPO / KTO / ORPO / SimPO / IPO / PPO / RewardModel / Pretrain / Embedding / BCO can now load pre-quantized checkpoints (TheBloke/Llama-2-7B-GPTQ, etc.) and train LoRA adapters on top. The compatibility-matrix cross-validators (HQQ × ZeRO-3 reject, EETQ × FSDP reject, AQLM × FSDP reject) already gate at config-load and apply task-agnostic.
- PPO reward model loads quantized too —
_load_reward_modelaccepts an optionaltcfgkwarg and routes the reward checkpoint through the same Quant Menu loader as the policy. A GPTQ-policy + GPTQ-reward PPO run no longer silently OOMs in fp16. - Defence-in-depth on
reward_modelfield —TrainingConfig.reward_modelfield validator now rejects null bytes and caps length at 512 chars at config-load, matching the policy already applied tocfg.base. - +131 net new tests (4930 → 5061) across the schema-gate widening matrix (11 tasks × 7 formats), MLX rejection regression per task, source-level invariants asserting each non-SFT trainer no longer carries the legacy
BitsAndBytesConfig(load_in_4bit=True ...)literal, and a live mock-based dispatch test for_load_reward_model.
Install / Upgrade
pip install --upgrade soup-cliSecurity
TrainingConfig.reward_modelfield validator: null-byte rejection + 512-char cap at config-load (defence-in-depth ahead of the runtime null-byte check inquant_menu._check_local_marker).- All 11 non-SFT trainers now route quantized loads through
build_quantization_config_for_loader, inheriting the v0.38.0 hardening (GPTQquantize_config.jsonprobe, AWQquant_config.jsonprobe, HQQnbitsallowlist, AQLM 2-bit lock, EETQ 8-bit lock, MXFP4 BNB validation, FP8 dequant). prepare_model_for_kbit_trainingtuple widened to includemxfp4so the BNB MXFP4 path correctly runs through kbit-prep instead of silently falling through.
Known Limitations
- Vision / audio modality + Quant Menu still rejected by the modality gate. The vision and audio paths in
sft.pyretain inlineBitsAndBytesConfigblocks because they need vision-specific kwargs (LLaVA processor, etc.) that the unified loader does not yet thread. Multi-modal Quant Menu wiring is a follow-up item. - Autopilot's quantization picker still recommends only
4bit/8bit/none. Now that Quant Menu is wired across all transformer trainers, a future patch can teach Autopilot to recommendgptq/awqwhenbasealready points at a pre-quantized checkpoint. tcfg.reward_modelis null-byte and length-validated at schema load but not path-containment-checked (is_under_cwd). This is consistent with howcfg.baseis treated — both can be HF repo IDs or absolute local paths. The Quant Menu loader's read-onlyos.path.isfileprobe is the only filesystem touch and usesos.path.basenamein error messages so absolute paths are not leaked.