github MakazhanAlpamys/Soup v0.75.0
v0.75.0 — MLX honours its config; unknown config keys refuse the load

latest release: v0.75.1
18 days ago

All 60 pull requests in this release came from outside the maintainer, by 22 people. What they found is the release: on the MLX backend, six training options were validated, documented, accepted — and read by nothing, so the same soup.yaml trained a different recipe on Apple Silicon than on a CUDA box, silently. And the deadline v0.74.0 promised is honoured: an unknown config key now refuses the load.

What's New

  • Breaking: an unknown config key refuses the load. v0.74.0 reported every key no config model declares — a typo like quantizaton, or a field that only exists on a newer Soup — as a warning that named v0.75 as the release that would start refusing. This is that release: soup train exits 1 before the training stack is imported, and the API / Web UI loader raises ValueError with the same text, naming the field you probably meant. Nothing is defaulted or substituted. The refusal says Refused. rather than Not applied.; the soup sweep guard, which always refused, changes wording the same way. The detector sees a config the way SoupConfig does: a root-level lora: block (the LlamaFactory / Axolotl spelling the schema has accepted and moved under training since v0.40.1) is remapped through the same function before the scan, so a spelling the validator accepts is never one the detector refuses — the release review caught the first draft refusing it. The two soup fetch examples configs that used it (llama-3.1-8b-lora, qwen2.5-7b-dpo) moved to the canonical training.lora; all 167 recipes, 21 templates and examples/configs/ scan clean. Key names are escaped before they reach the terminal (Rich markup and raw ESC bytes in a YAML key could restyle or spoof it), the scan stops at 100 findings and says so, and the Web UI's /api/train/start returns the key and the suggestion instead of a generic error. A config that must stay loadable on v0.74 as well needs the key removed, not renamed. (#627, #879)
  • MLX honours the config it accepted. data.train_on_responses_only (#683 by @Srinivasan8888 in #733), warmup_ratio / scheduler / weight_decay / optimizer (#686 in #734), training.max_grad_norm (#749 in #750), gradient_accumulation_steps (#684 by @AmirF194 in #696) and gradient_checkpointing (#685 in #698) were each validated and then dropped on backend: mlx. Only 8 of the 32 optimizer names have an MLX equivalent; the other 24 are refused by name instead of silently becoming AdamW. The response-only mask is Soup's own per-token mask, not mlx-lm's flag — upstream's masks a single prefix and supervises only the last assistant turn on multi-turn chat (measured: 772 supervised tokens unmasked, 71 with upstream's flag, 146 with a correct mask). MLX also drives the live dashboard, the SQLite tracker and the soup ui stream (#23 in #665), and reports validation loss (#739).
  • Validation loss existed nowhere. It was computed on every backend and thrown away: on_log read the training loss and never the eval one, and there was no metrics column and no event field to put it in. It is now recorded, streamed and shown on the panel (#23 by @Srinivasan8888 in #713).
  • Two guards for the pattern this release is named for. A declared config field that reaches no consumer now fails the suite (#748 by @Srinivasan8888 in #751; #770 records the guard's measured leak, so it is a ratchet, not a proof). And soup doctor --config soup.yaml lists the settings your task/backend does not read, from a declared table — inference over the import graph detected none of five independently-known MLX gaps, which is why the table is written by hand and guarded in both directions (#755 in #756). Scope is deliberately narrow: task=sft on backend=mlx, seven entries.
  • Breaking: grpo_variant: gspo is the published sequence-level objective (Qwen Team, arXiv:2507.18071; #723 by @kok-o in #744), replacing a column-centering heuristic in which a padding token also shifted the loss and gradient of every unmasked row sharing its column (#735 by @AmirF194). Existing gspo configs will not reproduce prior runs.
  • training.loraplus_lr_ratio crashed every run that set it. It was forwarded to TrainingArguments, which has no such field. It now builds a real PEFT LoRA+ optimizer (#738 by @abdulwaarith0), and optimizer + scheduler state is proven to survive save/resume (#724 by @SID-6921 in #747).
  • torch>=2.6.0 closes v0.74.0's headline known limitation: at torch 2.5.1, trl>=0.29 could not import and every preference trainer was dead. CI now proves the floor (#651 by @Samearth17 in #717).
  • Data pipeline agreement. soup data validate runs the converters so it agrees with the loader (#712 by @abdulwaarith0) and fails when no row is usable, with an optional minimum-valid-fraction threshold for CI (#811 by @akkupratap323 in #858); a non-dict message drops the row instead of aborting the load, on the multimodal (#670 by @abdulwaarith0) and chatml / audio / video converters (#676 by @BetterAndBetterII in #678); data.interleave no longer places a row on both sides of val_split (#701 by @AmirF194); data.streaming is honoured for a single Hub dataset name (#689 by @jagadeepmamidi in #710); the single-pass validator (#771 by @iam-saiteja).
  • Inference prompts are encoded the way training encoded them. soup chat, soup serve and soup infer re-added special tokens to an already-rendered chat template, so tuned models were prompted with a doubled BOS they never saw in training (#781 by @Konuktor in #782).
  • ULD cross-tokenizer distillation refuses wasserstein / topk_align on tokenizers that disagree instead of clamping ids into range and training on the teacher's logits for the wrong text with a finite, plausible loss (#704 by @AmirF194), and respects the response-only label mask and causal shift (#711).
  • Multipack FFD placement is O(N log N) via a segment tree, with the packing proven unchanged over 6,000 randomized cases — 96x faster at 30,000 rows (#694 by @dchaudhari7177 in #726).
  • Also: packing: true on TRL 0.29 (#691 by @jagadeepmamidi in #709); lora.r: 0 full fine-tuning on task: embedding (#700 by @kok-o in #705; #690 by @jagadeepmamidi in #708); PPO forwards training.epochs and ppo_kl_penalty to trl 0.29 (#721 by @be-student in #737); soup eval auto no longer dies with a TypeError after a successful eval (#752 in #754); soup ingest --source langfuse --pull fetches generations live (#204 by @Konuktor in #859); AWQ export requires explicit calibration data (#592 by @Faisal01011); soup doctor reports the [train] extra as one optional group (#854 by @wangzhengzhuo05); terminal plots honour NO_COLOR (#860 by @akkupratap323); soup migrate no longer refuses a valid config for being named *.jsonl (#675 by @swalla02); DeepSpeed's empty-param-group guard receives the model it inspects (#784 by @DYNOSuprovo in #786); the GEMM throughput forecast measures in the card's resolved stream dtype (#617 by @Samearth17 in #648); MLX recipe repo IDs repaired (#661 by @kok-o in #666); soup data validate rejects an unknown --format instead of reporting rows valid for it (#866 by @SID-6921 in #869); four recipes — deepseek-v4-flash-dpo (#662), kimi-k2.6-dpo (#664), qwen3.5-0.8b-grpo and qwen3.5-2b-grpo (#848 by @SID-6921 in #864) — taking the catalog to 167; a Turkish README behind a gate that checks what a translation was synced from (#852 by @Ercaner1988); two repo-wide AST ratchets on path containment (#775 by @Ercaner1988 in #778, #783 by @AaronProbha18 in #789); opt-in hardware-only telemetry, off unless SOUP_TELEMETRY=1 (#318 by @kok-o in #529).

Full list: CHANGELOG.md.

Install / Upgrade

pipx install soup-cli            # or: pip install -U soup-cli
pip install -U "soup-cli[train]" # the training stack: torch>=2.6.0, transformers>=5.16.1, trl>=0.29, peft>=0.20

Python 3.10–3.12. If a config of yours has been printing unknown config key warnings since v0.74.0, fix or remove those keys before upgrading — v0.75.0 refuses them.

Security

  • Web UI read endpoints and SSE streams require authentication (#687 by @kok-o in #707). Run configurations, logs and system metrics no longer answer unauthenticated. SSE authenticates via short-lived, single-use tickets exchanged over an authenticated POST, so a durable token never rides in a query string. A non-loopback bind without a valid token is rejected with exit 2.
  • soup ui --public no longer serves the FastAPI docs to the LAN (#731 by @Srinivasan8888 in #732). /openapi.json, /docs, /docs/oauth2-redirect and /redoc are absent (404) on a non-loopback bind and unchanged on loopback — removed rather than gated, because /docs is a browser navigation that cannot carry a Bearer header. Reconnaissance, not disclosure: every endpoint the schema described already answered 401 after #707.
  • The Web UI training subprocess cannot hang when no client reads its output — stdout is drained by a bounded background buffer (#688 by @kok-o in #706).
  • Telemetry is opt-in and hardware-only; SECURITY.md documents the egress policy. Supported version for security fixes: 0.75.x.

Known Limitations

  1. The unknown-key refusal is a hard break for a config that was silently carrying an unknown key. v0.74 warned and ignored it; v0.75 refuses it; neither applies it. A config that must load on both needs the key removed.
  2. soup doctor --config knows seven settings on one task/backend pair (sft / mlx) and reports nothing elsewhere rather than guessing. The field-consumer guard (#748) has a measured leak (#770): it catches a field nothing reads, not a field something reads wrongly. Its all-clear message overclaims — CUDA-only kernels such as use_liger are unread on MLX and absent from the table (#903).
  3. soup plan / soup apply do not run the unknown-key check — they read the YAML as a plain mapping (pre-existing; #894). soup draft distill / soup shrink surface a loader ValueError as a raw traceback (#895). Six command modules still carry a private copy of the terminal-escaping helper the loader now imports from soup_cli.utils.terminal (#896). The Web UI's YAML endpoints have no request-size cap (#897). The autodistill worker's receipt-PID check fails under a Windows venv launcher (#898).
  4. Streamed-NF4 bit-exactness does not hold on Blackwell (sm_120). Two CUDA-only tests fail on an RTX 5070 at 4.9e-4 / 3.9e-3 while the quantization: none variants pass, so the divergence is in the bitsandbytes NF4 path on that card, not in layer streaming. CI has no GPU runner and cannot see it. Root cause not established — #776.
  5. The Turkish README was re-synced for this release by the maintainer's release tooling, not by its translation owner. It passes the structural gate; @Ercaner1988's polish is the standard it was merged under.
  6. #371 stays open — the reward-hack mitigation controller has never had a valid pid_lagrangian run.
  7. Layer streaming remains BETA.

Measurement record

  • Gate record: three contributor records landed in this window and are indexed in benchmarks/README.md — the failed QuEST W4A4 SFT gate, published as written (#674 by @Shutaru in #856), pinned-store accounting for layer streaming (#655 by @tristangrech in #787), and an 8 GB M1 MLX SFT run with its harness (#663 by @Srinivasan8888). No maintainer gate: nothing in this release was gated by a new measurement.
  • Preprint (DOI 10.5281/zenodo.21771064): no measured number moves and the scope does not change — nothing in this window touches layer streaming's mechanism or its architecture list; the paper stays scoped to v0.73.0 as its header says.

Contributors

60 pull requests, 22 people, none of them the maintainer:

  • @Srinivasan8888 (15) — the MLX backend's dropped options (#733 response-only mask, #734 optimizer/schedule, #750 max_grad_norm), MLX driving the live display/tracker/SSE (#665) and reporting validation loss (#739), validation loss persisted and displayed at all (#713), the field-consumer guard (#751) and its measured leak (#770), soup doctor --config (#756), the --public docs-route removal (#732), the typer OptionInfo leak (#754), scheduled recipe model-id resolution (#715), the M1 MLX run record (#663), the smoke-adapter loadability assertion (#668), the deepseek-v4-flash-dpo recipe (#662).
  • @kok-o (7) — sequence-level GSPO (#744), Web UI auth on reads and SSE (#707), the subprocess drain (#706), lora.r: 0 on task: embedding (#705), opt-in telemetry (#529), the kimi-k2.6-dpo recipe (#664), MLX recipe repo IDs (#666).
  • @AmirF194 (7) — MLX gradient_accumulation_steps (#696) and gradient_checkpointing (#698), the gspo masked-token centring (#735), ULD refusing mismatched tokenizers (#704) and honouring the label mask (#711), interleave vs val_split (#701), the accelerate floor pin in soup doctor (#753).
  • @abdulwaarith0 (3) — LoRA+ crashing every run that set it (#738), soup data validate agreeing with the loader (#712), the multimodal drop path (#670).
  • @jagadeepmamidi (3) — packing: true on TRL 0.29 (#709), Hub data.streaming (#710), the embedding trainer's missing attributes (#708).
  • @Samearth17 (2) — the torch 2.6.0 floor that closes #651 (#717), the GEMM forecast dtype (#648).
  • @swalla02 (2) — the pipx / PEP 668 install lead (#673), soup migrate JSONL sniffing by content (#675).
  • @Konuktor (2) — the doubled-BOS inference encoding (#782), soup ingest --source langfuse --pull (#859).
  • @Ercaner1988 (2) — the Turkish README and its sync ratchet (#852), the repo-wide path-containment ratchet (#778).
  • @Shutaru (2) — the failed QuEST W4A4 gate, published as written (#856), the MLX harness display/tracker bridge (#703).
  • @akkupratap323 (2) — soup data validate exit codes and the minimum-valid-fraction gate (#858), terminal plots through Rich under NO_COLOR (#860).
  • @dchaudhari7177 — the multipack segment tree (#726).
  • @SID-6921 (3) — LoRA+ optimizer and scheduler state surviving save/resume (#747), the qwen3.5-0.8b-grpo / qwen3.5-2b-grpo recipes (#864), soup data validate refusing an unknown --format (#869).
  • @be-student — PPO forwarding the epoch budget and KL coefficient to trl 0.29 (#737).
  • @iam-saiteja — the single-pass dataset validator (#771).
  • @BetterAndBetterII — non-dict messages dropped by the chatml / audio / video converters (#678).
  • @Faisal01011 — explicit AWQ calibration data (#592).
  • @DYNOSuprovo — the DeepSpeed empty-param-group guard's missing model argument (#786).
  • @tristangrech — the pinned-store accounting record for #655 (#787).
  • @AaronProbha18 — extending the containment ratchet to contextlib.suppress (#789).
  • @wangzhengzhuo05 — soup doctor reporting the [train] extra as one optional group (#854).
  • @Amix29 — the Web UI front-door section in the README (#728).

Thank you. Every one of these is a defect a user would have hit first.

Don't miss a new Soup release

NewReleases is sending notifications on new releases.