What's New
v0.53.1 lifts six v0.53.0 deferred stubs from NotImplementedError to live wiring.
#82 — Autopilot detects pre-quantized bases
detect_prequantized_format(name, hf_config=None) and detect_prequantized_format_from_path(model_dir) in soup_cli/autopilot/decisions.py. Name-regex matrix over gptq / awq / aqlm / eetq / fp8 / mxfp4 with word-boundary anchoring (rejects agptqa), HQQ pattern with explicit bit extraction (hqq-4bit → hqq:4bit), config-probe via quantization_config.quant_method (canonical + bitsandbytes_4bit / bnb_4bit / nf4 → 4bit aliases). decide_quantization gains optional prequantized: Optional[str] kwarg that short-circuits the VRAM heuristic. Autopilot pipeline auto-applies: TheBloke/Llama-2-7B-Chat-GPTQ is now recommended gptq instead of stacking 4bit on top.
#142 — soup merge --save-format + soup export --format torchao
soup merge --save-format {fp16|4bit|4bit_forced} writes a single BNB-4bit-quantized merged checkpoint without the dequant→merge→requant cycle. soup export --format torchao --quant-config <yaml> runs torchao.quantize_ + save_pretrained with a per-scheme closed kwarg allowlist (Int4WeightOnly accepts {group_size, inner_k_tiles}, NVFP4 accepts nothing extra). load_quant_config enforces yaml.safe_load + 256 KB cap + extension allowlist + cwd containment + symlink rejection.
#139 — soup export --format gguf-ud --gguf-flavour <UD-Q4_K_XL | IQ2_M | Q4_0_4_4 | …>
Three-stage pipeline: HF → f16 GGUF → optional importance-matrix (UD ladder + low-bit IQ) → quantize. All subprocess invocations use argv-list form (no shell) + 30-min timeout. Calibration JSONL is sanitised (null-byte stripped, newlines collapsed to spaces, 8 KB per-line + 50 MB total cap) before being passed to llama.cpp imatrix. POSIX O_NOFOLLOW defeats the TOCTOU race between dispatch-time symlink check and open(). UD- prefix stripped before passing to llama-quantize (UD-Q4_K_XL → Q4_K_XL).
#109 — soup deploy autopilot --measure --tasks <jsonl>
Live Quant-Lobotomy scorecard. Loops every candidate quant through the v0.26.0 eval/quant_check scorer, renders OK/MINOR/MAJOR table, picks the best-by-delta candidate (matches the v0.33.0 #54 design intent). Results cached at ~/.soup/deploy_autopilot_cache.json keyed on (base, profile, eval-tasks) (atomic write, 0o600 perms on POSIX, symlink-rejected on both load AND save). Override via SOUP_DEPLOY_AUTOPILOT_CACHE confined to home / cwd / tempdir + null-byte / control-char rejection.
#70 / #72 — Manual QA log scripted
tests/qa/v053_qa.md ships exact reproduction recipes + acceptance criteria for the CUDA + llama.cpp smokes that can't run on CI runners.
Install / Upgrade
pip install -U soup-cliOr via Docker (GHCR):
docker pull ghcr.io/makazhanalpamys/soup:v0.53.1Security
- Per-scheme kwarg allowlist on
export_torchaorejects dunder keys + unknown params before the splat into the torchao factory (defeats YAML-driven kwarg injection). enforce_under_cwd_and_no_symlinkconsolidated inutils/paths.py— single source of truth for the v0.33.0 #22 TOCTOU pattern; reused by merge / export / advanced-GGUF dispatch.detect_prequantized_format_from_pathcwd-contained + symlink-rejected on<model_dir>/config.json(soft-probe: out-of-cwd paths silently returnNoneso HF Hub repo IDs aren't rejected).- POSIX
O_NOFOLLOWon the calibration-data read closes the TOCTOU race between dispatch-time check andopen(). _safe_stderrRich-escapes subprocess stderr before embedding inRuntimeErrormessages so a crafted llama.cpp error cannot inject markup into the operator-facing panel.- Realpath-verified
convert_hf_to_gguf.pystays insidellama_cpp_dirafter symlink resolution — defends against a symlinked script escape. - Atomic cache write via
tempfile.mkstemp+os.replace, 0o600 perms on POSIX, symlink rejection on BOTHload_cacheandsave_cache(defence-in-depth).
Four review agents (python / code / security / tdd) ran with every CRITICAL / HIGH / MEDIUM / LOW finding fixed before tagging.
Known Limitations
#70GGUF and#72AWQ/GPTQ manual QA smokes remain PENDING (require a CUDA box + llama.cpp build; recipes scripted intests/qa/v053_qa.md)._DEPLOY_MEASURE_BEFORE_GEN/_DEPLOY_MEASURE_AFTER_FACTORYmodule-level injection is a stop-gap until v0.46.1 ships first-party transformers / vLLM generator factories. NOT a public API.- CPU-only smoke for
merge_4bit/export_torchao— code-side is fully exercised via mocks; real BNB-4bit + torchao kernels need CUDA + extra deps. _prepare_calibration_textaccepts JSONL withtext/prompt/contentaliases + raw text fallback; other formats (parquet, markdown) are out of scope — convert upstream.#109cache key truncatesbase_shato 16 hex chars at the call site (collision probability ≈ 1 in 2³² across ~4 billion entries).#82pre-quantized detection is heuristic — name regex + localconfig.jsonprobe. HF Hub repo IDs without local download fall back to name-only matching.enforce_under_cwd_and_no_symlinkchecks only the leaf path — internal symlinks inside a contained tree are not detected; per-file leaf check at each site is the threat-model boundary.
Stats
- 7610 → 7722 tests (+112 across 4 new files:
test_v0531_82.py,test_v0531_109.py,test_v0531_139.py,test_v0531_142.py) - 180 → 184 test files
- New module:
soup_cli/utils/deploy_measure.py - Manual QA log:
tests/qa/v053_qa.md
Co-Authored-By: Claude Opus 4.7 (1M context) noreply@anthropic.com