github MakazhanAlpamys/Soup v0.40.3
v0.40.3 — Stub-to-live wave 1 (#33, #64)

latest releases: v0.75.2, v0.75.1, v0.75.0...
4 months ago

What's New

Three v0.X.0 deferred-stub features become live runtime — closes #33 (harvester judge filter + serve trace log) and #64 (live CUDA OOM probe). #65 (multipack live wiring in HF Trainer) remains deferred to v0.40.4 after the adversarial 5th-review pass surfaced a Sampler[int] vs list[list[int]] shape mismatch with HF Trainer's DataLoader.

  • Live CUDA batch-size probe — auto_batch_size_strategy: probe now runs ONE forward+backward+step on a synthetic batch per candidate before training. On torch.cuda.OutOfMemoryError the probe halves; otherwise it doubles. Result is cached per (model, max_length, quant, lora_r, gpu) tuple so the next run short-circuits. SFT-only this release.
  • Multipack helpers — make_multipack_trainer_class (lru-cached subclass factory), attach_multipack_state (with bool-rejection + empty-lengths reject), lengths_from_dataset (with all-zero WARNING), detect_arch_name. The schema gate already rejects multipack: true on unsupported tasks. Live trainer wiring is deferred to v0.40.4 — multipack: true currently prints a yellow advisory and falls back to the standard sampler.
  • soup data from-traces --judge — optional LLM-as-a-judge pass over harvested preference pairs. --judge-provider openai|server|ollama, --judge-model gpt-4o-mini, --min-confidence 0.7. Drops pairs whose normalised (chosen - rejected) confidence falls below threshold. Per-pair backend exceptions counted (not crashed); lazy itertools.islice cap; cost-shock warning before the loop (2× per pair).
  • soup serve --trace-log <path> — passive append-only JSONL request log ({prompt, response, latency_ms, tokens, ts} per chat completion). Path-containment validated, 100 MB rotation cap (one backup retained, symlink-reject on rotate), hf_* / sk-* / Bearer … token shapes redacted to <redacted> before write (recursively in extra dict too). Streaming SSE path also records.

Behaviour change: v0.40.2 users with auto_batch_size_strategy: probe were silently getting the static fallback. v0.40.3 actually runs a CUDA probe on first run (~5–30s, cached).

Install / Upgrade

pip install --upgrade soup-cli

Security

  • _redact_secrets regex with Bearer body excluding . so end-of-sentence period survives.
  • Symlink-reject on <trace-log>.1 rotation backup (TOCTOU defence, mirrors v0.33.0 #22 policy).
  • judge_provider allowlist validation at the CLI boundary BEFORE constructor invocation.
  • len(tokenizer) (not vocab_size) bounds pad_id in CUDA probe — prevents pad folding to a random byte token on extended-vocab tokenizers (Llama-3 + <|pad|>).
  • Path containment via shared is_under_cwd on --trace-log path.

Known Limitations

  1. Multipack live HF Trainer wiring deferred to v0.40.4 — adversarial review of v0.40.3 caught a Sampler[int] vs list[list[int]] mismatch; setting multipack: true prints a yellow advisory and falls back to the standard sampler. Live wiring requires a get_train_dataloader override.
  2. Live CUDA probe (make_cuda_probe_fn) is wired in SFT only; non-SFT trainers still use the static estimate.
  3. TraceLogWriter retains exactly ONE backup file (<path>.1); operators wanting longer retention should arrange external rotation.
  4. Multi-worker uvicorn deployments writing to the same --trace-log path race on rotation — single-process assumption.
  5. Custom HF Space templates from v0.40.2 still always create the Space with space_sdk="gradio"; tracked for v0.40.4+.
  6. --judge issues TWO judge calls per pair (chosen + rejected); a yellow projected-call-count warning prints before the loop.

Don't miss a new Soup release

NewReleases is sending notifications on new releases.