Highlights
779 PRs from 227 contributors.
New models
| Model | Type | Cookbook |
|---|---|---|
| DeepSeek-V4.1 Flash | LLM / VLM | link |
| GigaChat 3.5 | LLM / VLM | link |
| IQuest-Q1 | LLM / VLM | link |
| MiMo-V2.6 / MiMo-V2.6-Pro | LLM / VLM | link |
| Ling-3.0-flash-VL | LLM / VLM | link |
| DiffusionGemma | Diffusion | link |
| Qwen-Image 2.1 | Diffusion | link |
| Anima Base v1.0 | Diffusion | link |
| Ming-Image 0.1 Design / Design-Layer | Diffusion | link |
| FLUX 3 Action | Diffusion | link |
Key features
- PD instances can switch between prefill and decode on the fly, no restart needed (#28403).
- The prefix cache now runs on a Rust core by default (#39627).
- DeepSeek-V4.1 gets 22% faster first token on long prompts (#40352).
- Kimi K3 gets 20.6% higher prefill throughput in PD serving (#40045).
- More accurate results under PP, DP attention, and CP, with layer communication now handled by SGLang (guide, #41557).
- New Decisions API (
/v1/decisions) turns an LLM / VLM into a low-latency classifier and scorer (docs, #41208). - Score API (
/v1/score) scores all candidates in one request (#38965, #41188). - Run MiniMax-H3 with SGLang Diffusion inside ComfyUI (#35990).
To upgrade
uv pip install --prerelease=allow sglang==0.5.21| Platform | Docker image |
|---|---|
| NVIDIA (CUDA 13) | lmsysorg/sglang:v0.5.21
|
| AMD MI35x | lmsysorg/sglang:v0.5.21-rocm10-mi35x
|
| AMD MI30x | lmsysorg/sglang:v0.5.21-rocm10-mi30x
|
| Intel GPU | lmsysorg/sglang:v0.5.21-xpu
|
| Intel CPU | lmsysorg/sglang:v0.5.21-xeon
|
Full Release Notes
Speculative Decoding
- Pipeline parallelism x speculative decoding (EAGLE/MTP) compatibility: #30775
- Support XQA backend for SpecDec verify: #32269
- [Spec] Windowed draft-decode attention for built-in EAGLE / MTP drafts: #32673
- Avoid materializing GDN QKV tensors during target verification: #33778
- [Spec] Fix CDF boundary handling in
TreeSpeculativeSamplingTargetOnly: #35798 - [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts: #37462
- Add out-of-tree DFlash extension points: #38740
- [Spec] Add explicit prefill shared-read capability for plugins: #39502
- [Fix] Don't write conv state from the fused KDA verify kernel: #39524
- [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism: #39643
- Use runtime token widths for Triton speculative verification: #39859
- avoid host sync in DSpark prefill slot expansion: #40111
- [Experimental] Preserve speculative decoding during prefill across DP ranks: #40118
- [Spec][PP] Launch extend microbatches before the spec output exchange: #40499
- [KDA] Enable ReplaySSM for GLM-5.3 Flash: #40517
- [LFM2-VL] Add DSpark speculative decoding (1.66x to 2.56x speedup at batch 1): #40651
- [Spec] Support DFLASH for Kimi K3: #40794
- [Fix] Reduce Nemotron MTP attention outputs once: #40800
- [Fix] Capture complete Nemotron auxiliary hidden states: #40801
- [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune: #41138
- Fix mixed chunk prefill with DP speculative coordination: #41179
- [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft: #41194
Piecewise & Breakable CUDA Graph
- [Fix] Preserve model runner contracts in prefill CUDA graphs: #35452
- [DSV4][BCG] Optimize the heavy memory use of C4 Indexer when BCG is enabled (capture memory down 58% to 18 GB): #36534
- [Fix] Don't free the multi-CTAs KV counter the decode graphs captured: #39175
- [CPU] Avoid prefill CP predicates during decode graph capture: #39690
- [Runtime] Add decode CUDA graph hooks for eager logits processing: #40222
- [Kimi K3] Fix CUDA graph stream explosion: #40640
- [DSpark] Fix draft CUDA graph stream explosion: #40658
- [Fix] Give the full prefill CUDA graph replay view the captured bucket's input_ids: #40851
- [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention: #41311
Attention Backends
- Allow attention layers to opt out of the prefill wrapper: #40683
- [Fix] Use cached prefix lengths for FlashInfer full-attention ragged prefill: #40796
MoE & Expert Parallelism
- [3/N] elastic-ep: Recapture decode CUDA graphs after scale-up: #33723
- Fix int4 MoE tuner config filename: #35260
- [MegaMoE] Wire Qwen MoE blocks to DeepGEMM MegaMoE (MXFP4 and NVFP4 experts): #38080
- [Feature] Support BF16 and batch-invariant inference with DeepEP v2: #38160
- [Perf] Optimize w4a8 MoE for glm5.2 on H200: #38220
- [Fix] Guard FlashInfer CUTLASS MoE against 0-token inputs: #38780
- [Kernel] Add H20 block-FP8 MoE configs for GLM-5.3-Flash EP4/EP8: #38913
- Accept MXFP8 dispatch in FlashInfer A2A TRT-LLM MoE: #39613
- [MoE] Honor swiglu_limit clamped activation in flashinfer_cutlass runner: #39939
- Support MXFP8 and deferred route weighting in DeepEP v2: #40030
- [MoE] Disable FlashInfer fused finalize by default for numerical accuracy: #40105
- [Fix] Fix top-1 MoE routing with non-unit scaling: #40187
- [DeepEP v2] support GLM-5.3-Flash (Glm5NextForConditionalGeneration): #40466
- [Fix] Decide the MoE padded-row bound from the layer scatter mode: #40672
- [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound: #41201
Quantization
- [MXFP8-KV] Skip writes to the reserved CUDA-graph padding slot: #35351
- feat(kv-cache): support SM100 NVFP4 GenMHA and speculative decoding (1.355x faster decode attention kernel at 1M context): #36340
- [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size: #38726
- fix(modelopt): dispatch NVFP4 MoE on the cached backend, not the live global: #38932
- [Quant] Serve 32-wide-K ue8m0 block-FP8 linears through the FlashInfer MXFP8 GEMMs: #40039
- [ModelOpt][PP] Keep BF16 shared experts out of the NVFP4 fusion so TP1 pipeline stages can load: #40628
Parallelism & Disaggregation
- [PD] Introduce runtime role switching between prefill and decode: #28403
- disaggregation: Fix FakeKVSender queue accumulation: #28652
- [PD] Keep EAGLE DP graph and token metadata consistent: #32196
- [NVIDIA][comm] Merge EP+MoE-TP post-experts all-reduces into one _TP reduction: #32963
- Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++: #37442
- [KV-Shard 2/4] Sharded pools: #37615
- [KV-Shard 3/4] Enable Control plane: #38468
- [KV-Shard 3/4] Enable Control Plane B: #39964
- Fix device context during NIXL backend initialization: #38774
- [PD] Do not admit intake-rejected requests to a PD handoff: #38935
- Reduce decode bootstrap latency with request-owned speculative KV (mean TTFT 19.4 to 15.1 s on DeepSeek-V4-Pro 1P1D, MI355X, concurrency 256): #38978
- Fix disagg PP MTP for GLM-5.2: #39378
- [PD] Add optional KV transfer checksums: #39500
- [DCP] Use logical token capacity for PD admission and load reporting: #39731
- Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.): #39816
- [PD] Validate Mooncake EFA allocator compatibility: #39973
- [PD] Skip singleton transfer-status all-reduces: #40003
- [PD] Enable optimistic prefill with buffer-only L3 write-through HiCache: #40043
- [Kimi K3][PD] support pp prefill + dcp decode with dspark: #40045
- [PD] Add decode host receive for custom transfer backends: #40238
- [PD] Allow decode radix cache and HiCache L1/L2 with DCP: #40263
- [PD] Enter the custom mem pool once when allocating DCP pack buffers: #40284
- [PD] Bound cached-prefix DCP transfers by pack capacity (cached-prefix throughput about 10.9x higher at concurrency 8 on Kimi Linear, 8x B200): #40376
- [PD] Pack draft KV head slices for DCP transfers: #40500
- [PD] Preserve abort ACKs until in-flight KV transfers drain: #40645
- Fix NIXL transfer of MXFP8 KV block scales: #40792
- [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM: #40793
- [DP Attention] Publish DP buffer sizes from a ForwardBatch: #40858
- [PD] Enable deferred decode-side KV release by default: #41023
- [PD] Add a
nonedecode retraction backup and subclass seams in the PD queues: #41103 - [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match: #41261
- [PD] Fan drain abort ACKs out to every decode peer of the room: #41402
- [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts: #41404
- [Fix] Stop deferring the last layer's FFN all-reduce in five models: #40868
- [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send: #41079
- [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture: #41080
- [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention: #41082
- [Fix] Broadcast requests along attention CP before attention TP: #41083
- [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch: #41193
- [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv: #41195
- [Fix] Keep one copy of CP-replicated rows in the DP gather: #41422
- [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum: #41423
- [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather: #41432
- [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism: #41433
- [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators: #41436
- [Refactor] Group communicator fusion and CP adapters: #41547
- [Refactor] Centralize decoder output access: #41548
- [Refactor] Carry residual state across stage boundaries: #41549
- [Refactor] Capture auxiliary states at residual reads: #41550
- [Refactor] Select reduction fusion at the consumer: #41551
- [Refactor] Construct independent decoder stage boundaries: #41552
- [Refactor] Migrate specialized decoder and overlap boundaries: #41553
- [Refactor] Retire the layer facade and simplify boundary internals: #41554
- [Refactor] Rename the module to layer_boundary: #41555
- [Refactor] Group layer boundary unit tests: #41556
- [Refactor] Document layer boundary contracts and integration: #41557
Scheduler & Runtime
- feat: support custom OTLP trace service name: #35802
- [RL] Release the weight-checker snapshot once compare passes: #37284
- [Runtime] Let out-of-tree platforms provide full graph backends: #37969
- Fix: post-load staging regression breaks offload meta/sharded_gpu modes: #38779
- feat: add kv hint envelope to request transport: #38891
- [Observability] Fix negative queue_time for retracted requests: #39312
- Fix first-token metadata and reused attention-layer indexing: #39328
- perf(engine): avoid timed waits for Engine responses (+37.7% output throughput in a Dynamo TCP push benchmark at the default batch_notify_size): #39486
- [Misc] Fix tool-call index, graph padded-row count, and prefill-graph input_embeds refresh: #39574
- [Misc] Merge FlashInfer autotune caches across spec workers, pad MXFP4 TP shards, drop dead ngram attrs: #39678
- [Observability] Expose python/rust frontend identity in
/server_info: #39993 - [Metrics] Propagate idle gaps across all scheduler loops: #40004
- [Scheduler] Count complete prefill bursts and their tokens: #40006
- [Scheduler] Add shortest-prefill-first scheduling: #40024
- [Profiler] Ignore PREBUILT batches in profile-by-stage: #40098
- Enable optimistic prefill for Mamba radix-cache models: #40184
- [Perf] Fork-safe import: no CUDA context at import time, lighter argument parsing: #40201
- [Feature] Support --tokenizer-worker-num > 1 in the offline Engine API: #40260
- Clean up startup logging and streamline log audits: #40526
- [Fix] Run KV canary hooks for context-parallel prefill: #40642
- [RL] Fix Kimi K3 expert-count lookup for routed-expert capture: #40700
- [RL] Add RL weight-update sessions and support updating spec draft runners: #40777
- [RL] Keep pause_generation and weight updates from deadlocking each other: #40779
- [Misc] Remove deprecated endpoints, env vars and aliases past two releases: #40795
- [Metrics] Log forward and forward+idle occupancy over total wall time: #40802
- [RL] Keep DSA cuda-graph state and the graph pool intact across TMS pause/resume: #40804
- [Fix] Recover from stale torch extension locks in every
cpp_extensionloader: #40989 - [Perf] Lazy-load built-in model definitions and nixl_ep at startup: #41061
- [Model Loader] Stop checkpoint prefetch after iterator completion: #41588
- [Benchmark] Limit warmup concurrency in serving benchmark: #39398
- Support LongBench v2 in one-batch server benchmarks: #39874
- [Benchmark] Take each request's prompt length from the server: #39889
- [Benchmark] Add agentic rollout simulator and offline explorer: #40034
- [Benchmark] Optionally clear HiCache storage between cases: #40659
Sampling
- fix(sampling): validate sampling_seed is an int within int64 range: #28960
- perf(sampling): avoid GPU syncs when applying custom logit processors: #39234
- Use pinned memory for asynchronous sampling metadata transfers: #39777
- [Logprob] Borrow graph-pool memory for input logprob logits construction: #40007
- [Logprob] Serve input-logprob temporaries from CUDA-graph-pool dead space: #40038
- [Sampling] Add selected/support sampling logprob modes: #40932
- [Sampling] Stream sampling masks as per-request arrays: #40986
- [PD] Keep the sampling mask of a replayed rebootstrap token: #41235
HiCache & Radix Cache
- [Radix Cache] Sync Rust TreeCore and make it the default: #39627
- Add Agentic-Aware Tail-Optimized LRU eviction to the unified radix cache (opt-in; p99 ITL down 43.9% at concurrency 32 on DeepSeek-V4-Pro FP4, 4x B300, agentic trace): #34012
- Feat: Add TensorCast storage as a new HiCache backend: #27265
- [HiCache] fix: resolve Mooncake local_hostname per node for runtime attach: #29668
- [Fix] fix(hicache): wait for decode offload before retraction: #30899
- [HiSparse] Add MHA hisparse support for MiniMax M3: #31446
- [Mooncake] Fix silent SSD offload corruption when TP/PP ranks share ssd_offload_path: #31926
- [Scheduler] Align
RadixCacheno-insert cleanup withkv_len_to_handle: #35204 - [PP + HiCache] Add PP Prefetch Tickets for eager cross-stage storage prefetch: #36700
- Use a shared byte budget for unified hybrid-SWA memory: #36729
- Gate mamba extra-buffer predicates on uses_mamba_radix_cache: #37474
- [Unified Memory] Hierarchical cache for every unified pool shape: #37507
- [HiCache] Fix sparse hybrid transfer layer IDs: #37870
- [Unified Memory] fix: preserve FP8 dtype in unified MHA pool: #38133
- [HiCache] Yield idle scheduler so storage workers can drain: #38504
- [LMCache] Support lmcache unified radix cache: #38652
- [Unified Cache] Dedup replicated MLA/DSA KV in the UMBP direct linker: #38778
- [Perf] Fuse SWA page lookup and mapping clear: #38948
- [HiCache][Perf] fix: batch HiCache D2H submits per step for hybrid pools: #39050
- [HiCache] Back up MXFP8 KV scales in the host pool: #39089
- [Mamba] Fix checkpoint depth for prefixes that end off the radix page: #39115
- [HiCache] Label radix-cache metrics per rank and split the "shrunk" prefetch reason: #39280
- bugfix:fix unifiedcache c128 radix cache management: #39426
- Support unified memory page-envelope transfers in PD: #39477
- Support unified memory decode host pools: #39478
- [Unified Cache][9/N] add opt-in MLA load deduplication for Mooncake Linker (TTFT 3.85 to 2.23 s with a 256K cached prefix, GLM-5.2 on 8x H20): #39565
- [HiCache] Forward prefix metadata to v2 storage calls: #39567
- [HiCache] Keep hybrid transfer layer maps stage-local under PP: #39699
- [HiCache] Add the page-unified KV load-back JIT kernel: #39726
- [2/N][Kernel] Fuse padding-preserving HiSparse slot translation: #39837
- [Unified Tree] fix: exempt host-locked aux nodes from the sanity_check host-LRU check: #39980
- [HiCache] Read the in-flight buffer backup's node id from its snapshot in sanity_check: #40013
- [HiCache] Stop arming a prefetch retry for a too-short storage span: #40042
- [MemCache] Release up to
owned_kv_lenon radix cache insert: #40075 - [HiCache] Auto-size the host pool to fit available host memory: #40135
- Preallocate HiCache MHA staging before post-capture KV sizing: #40256
- Fix prefetch attempt cleanup on abort: #40262
- [HiCache] TMA-staged host<->device KV transfer kernel (sm_90+): #40278
- [HiCache] Size MHA host pools from device row width: #40304
- [Fix] Missing SWA eviction during decode preallocation: #40309
- [HiCache] fix: bound the controller reset join so a stalled storage thread cannot hang the scheduler: #40312
- Remove swa and mamba radix cache: #40313
- [Fix] Forward SWA prealloc reclaim through the DSV4 HiSparse allocator: #40354
- [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain: #40456
- [HiCache] Make host reclamation independent of transfer order: #40512
- [HiCache] Demote internal-node mamba states on write_back eviction: #40680
- [HiCache] Demote internal-node SWA KV to host on write_back eviction instead of dropping it: #40712
- [MemCache] Remove the experimental C++ radix tree: #40775
- [MemCache] Clean up SWA/Mamba radix cache leftovers and drop SGLANG_ENABLE_UNIFIED_RADIX_TREE: #40780
- [HiCache] Remove the unused HiRadixCache: #40787
- [MemCache] Free the rows below the SWA evict floor on all-SWA request release: #40798
- [HiCache] Batch buffer-only KV backups within each flush: #40960
- [Fix] Derive per-runner hybrid SWA layer ids on ModelLayerInfo instead of mutating ModelConfig: #40983
- [HiCache] fix: Drain pending backups before internal Mamba write-back: #41092
- [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers: #41144
- Let predicate-registered linear-attention models carry the mamba radix-cache leaves: #41165
- [Unified Memory] Honor move gates in float relocation and size auto HiCache from host capacity: #41248
- [MemCache] Unify component eviction cursors and lock receipts: #41276
- [MemCache] Never free the protected prefix on request release: #41312
- Make sliding-window caching and speculative batch padding extensible: #41325
- [MemCache] Fix LMCache component cursors and per-cache backend selection: #41328
LoRA
- [LoRA] Support DP attention in LoRA backends: #36389
- [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard: #39379
Multimodal
- fix(multimodal): return 400 for corrupt image inputs: #28131
- [Multimodal] Avoid CUDA placement on non-CUDA platforms: #38750
- Fix Mistral3 retaining every vision-tower layer to read one: #39185
- perf(multimodal): offload CPU feature hashing with bounded admission: #39539
- Carry deferred attention operands and reuse multimodal shared memory: #39870
- [MM] Skip VMM error gathers for text-only requests: #40005
- [MM] Copy placeholder ids to CUDA asynchronously: #40010
- [MM] Keep scheduler padding in packed token arrays: #40357
- [Fix] Raise on undelivered embeddings in
send_with_url, fix broken tests: #40502 - Fix multimodal feature offload races: #40621
- [VLM] Introduce FA4 into ViT for SM100/SM103: #41344
Model Support & Optimizations
- [Feature] Gigachat 3.5 support: #29189
- [dLLM] Support DiffusionGemma serving: #34061
- Add Ling-3.0-flash-VL model support: #38526
- dsv4.1: remaining model and runtime integration: #38798
- [DSv4.1] Score prefill consumer index layers on candidate blocks with DeepGEMM: #40352
- [Feature] Xiaomi MiMo-V2.6/MiMo-V2.6-Pro day0 support: #40448
- [Model] Add IQuest Q1 support and MTP draft: #41590
- fix(gemma4): set lm_head_is_tied for Gemma4UnifiedForConditionalGeneration: #35809
- [Fix] Preserve YaRN scaling when extending rotary caches: #38786
- [Kimi-K3] O(1) expert weight lookup in load_weights (expert-name matching 21.45 to 0.13 s at load): #38805
- [Model] Serve DeepSeek-OCR-2 with its official 768px local-crop geometry: #38996
- Fix GLM-OCR MTP multimodal embeddings and positions: #39088
- [DSV4] Chunk the indexer MQA logits by query rows under a free-memory budget: #39095
- [Perf] Fuse the glm5_next mHC attn->MLP boundary: #39200
- [DSV4] fix: keep the TileLang JIT cache under SGLANG_CACHE_DIR: #39364
- Fuse GLM-5.3-Flash KDA projections and prefill metadata: #39688
- [GLM-5.3-Flash] Reduce KPool planning synchronization and overlap indexer preparation: #39695
- [Fix] Fix GLM5 mHC PP forward: #39720
- [DSV4] Generalize attention metadata, sparse prefill, and KV pool over compress ratios: #39921
- [Qwen3.8 Next] Fuse Qwen PLE gate and convolution preparation for target verify: #40041
- Support GLM-5.3-Flash hybrid attention CPU offload and PD index mapping: #40310
- [DSV4] fix: size the C4 state ring by the page it is addressed by: #40337
- [Fix] Keep mHC context out of non-V4 compiled MoE forwards: #40353
- [Qwen3.8-Next] Pipeline-parallel serving and PD-prefill MTP for Qwen4-Exp: #40501
- Fix GLM-5.3 forget-gate shape for nvCUTEDSL verify: #40607
- [Fix] Avoid duplicate residual in LongCat MoE shortcut: #40799
- [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget: #40854
- [DSV4] Budget the ratio-2 pair state pool in DSV4PoolConfigurator: #41048
- [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator: #41049
- [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B: #41059
- [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting: #41090
- [DSV4] Account for FlashMLA physical KV page padding in memory budgets: #41091
- [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping: #41164
- [Qwen3.8 Next] Fuse small CUDA graph input buffer copies: #41166
Kernel Library
- [Kernel] Share the warp vectorized copy and enforce its alignment: #36176
- fix(kda_prefill): fence shared writes before async proxy reads: #39124
- [Kernel] Coalesce the KDA CuTe DSL decode state transpose: ~3x faster, bit-identical: #39680
- [Kernel] Move CUDA and ROCm speculative kernels to JIT: #40033
- Fix TopK v2 fallback when 16-block cluster capacity is zero: #40163
- [Kernel] Fuse hc_combine_norm for mid-size verify batches (9-96 rows): #40208
- [KDA] Fix missing beta sigmoid in PTX prefill: #40685
- [JIT] Add an occupancy-preserving L1 carveout preference: #40767
- [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving): #41223
Configuration System
- Add registration for external model configurations: #39452
OpenAI-Compatible API
- [Fix] Guard conditional top-logprob keys in the completions echo path: #34776
- [Fix] Prevalidate JSON Schema support per grammar backend: #37839
- fix(openai): recover logprobs token bytes from token_id (UTF-8 fragments): #38604
- [Score API] Setwise Scoring Support: #38965
- feat: use XGrammar V4.1 DSML parameter constraints: #39026
- fix(function_call): buffer complete DeepSeek DSML invokes: #39632
- Fix corrupted chat prompts on mistral_common tokenizers (tool_choice auto never fires): #39773
- [Fix] Keep Inkling automatic tool grammar active across the response: #40468
- fix(grpc): expose native response timeout as a server argument: #40644
- [Feature] Add per-item candidate token scoring and calibration: #40826
- Add MiniMax arch fallback to auto parser resolution: #40930
- [Score API] Setwise scoring: CausalLM support (batched + --enable-mis): #41188
- [Feature] Add a System One compatible /v1/systemone route (also ships the
/v1/decisionsroute it builds on): #41208 - Fix chat template cache key order: #41517
Rust Server
The experimental Rust router and frontend (experimental/sgl-router, the Rust renderer) keep moving toward parity with the Python server.
- [Rust Renderer] Standalone preprocessing: #36718
- [SGL Router] Render chat prompts with dynamo-render: #38983
- [Router] Fleet-wide sampling contract 1/3: the config surface: #39000
- [Router] Fleet-wide sampling contract 2/3: enforce and inject per request: #39001
- [Router] Fleet-wide sampling contract 3/3: splice injection without re-serializing: #39002
- [Router] Resolve a wire protocol per worker at registration: #39004
- [Router] Speak cleartext h2c on both edges: serve it inbound, forward it outbound: #39006
- [Router] Count open HTTP exchanges until their response body finishes: #39014
- [Router] Add the shutdown-drain configuration surface: #39015
- [Router] Drain readiness before SIGTERM shutdown so k8s deregisters the pod first: #39016
- [Router] Real-GPU e2e coverage for storage-tier-aware cache routing (3/4): #39110
- [Router] Add a KV storage-tier coverage row to the Grafana dashboard (4/4): #39111
- [Router] Improve SGLang chat render parity: #39133
- [Router] Shard the cache-aware KV tree by chain root: #39167
- [Router] Add --worker-queue-limit: stop sending cache-affinity traffic to a queueing worker: #39168
- [Router] Pin to the prefix owner when the whole fleet is queueing (--saturation-queue-floor): #39169
- [Router] Sample k random candidates for the min-load fallback (--min-load-choices): #39170
- Add external multimodal processors to the Rust frontend: #39329
- [Rust] Extract a transport-neutral frontend core: #39385
- [SGL Router] Forward input_ids only for string content; count tokenize errors only when forwardable: #39458
- [Router] Abort the engine when a client disconnects mid-request: #39461
- [Router] Derive error status from a failure class; preserve the worker's status (1/3): #39463
- [Router] Treat an upstream 503/429 as backpressure, not a breaker fault (2/3): #39464
- [Router] Log every request at one site; derive its outcome from the final status (3/3): #39465
- [SGL Router] Share model-file discovery for chat formatters: #39485
- Rust server unify datapath for mm and generate requests: #39679
- [SGL Router] Add Kimi-K3 rendering with SGLang parity: #40390
- [SGL Router] Bound streaming lifetimes and release guards on idle disconnect: #40391
- Port chat_parsing core: #40477
- [SGL Router] Match DeepSeek V4 rendering to SGLang: #40530
- [SGL Router] Add SGLang-compatible DeepSeek V4.1 Flash rendering: #40532
- [SGL Router] Release cancelled circuit-breaker probes: #40603
- [SGL Router] Fix readiness, IPv6 discovery, logging, and model validation: #40604
- [Router] Give the cache-aware tree a snapshot surface (1/13): #40687
- [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13): #40688
- [Router] Name a replica's siblings with --kv-peer-selector (3/13): #40689
- [Rust Renderer] decouple renderer sampling from protocols: #40747
- [SGL Router] Launch reorg routing with existing policy options: #40766
- [SGL Router] Track input_ids forwarding outcomes per chat request: #41185
- [SGL Router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs: #41221
- [SGL Router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4: #41226
- [Rust Frontend] Decode input_ids without untagged buffering (input_ids deserialization about 45% faster): #41246
- [SGL Router] Book the input_ids forwarding outcome only for built bodies: #41342
Simulator
- [Simulator][Compatibility] Adapt to latest KV cache pool interfaces: #40418
SGLang-Diffusion
- [Diffusion] model: support qwen-image-2.1: #39983
- [Diffusion] Add native Anima Base v1.0 support: #41011
- [Diffusion] Support FLUX 3 Action robot policies: #41066
- [Diffusion] Add native Ming-Image Design and Design-Layer support: #41067
- [Diffusion] add /metrics support: #19084
- [Diffusion] Preserve per-sample rollout trajectories across multi-output merge: #34416
- [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager: #34417
- [Diffusion] Apply latent-ids and packing to caller-provided initial latents: #34418
- [Diffusion] MiniMax-H3 Spectrum skip-step + fused RMSNorm/AdaLN: #35684
- [Diffusion] Add MiniMax-H3 to ComfyUI integrated mode: #35990
- [Diffusion] In-place LoRA merge/unmerge under layerwise offload: #36192
- [Diffusion] feature: out of tree platform support: #37547
- [Diffusion] support multiple task types for pipelines: #38762
- [Diffusion] Enable shared RMSNorm dispatch for SenseNova-U1: #39705
- [Diffusion] Preserve explicit attention backends during autotune: #39882
- [Diffusion] Cache-DiT 1.5.1: DMD Calibrator, SVDQuant DQ, etc.: #40104
- [Diffusion][MiniMax-H3] Add SM120 Sage compute for SubBlock sparse attention: #40116
- [Diffusion] attention: add fp8_fa_sm120 FP8 backend for SM120 GPUs: #40175
- [Diffusion] Fuse lossless SenseNova RoPE for 5% faster H200 inference: #40374
- [Diffusion] Fuse rounded SwiGLU for quantized MiniMax-H3 MLPs (FastH3 FP8 worker E2E 1.9% lower on 2x H200): #40378
- [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG: #40384
- [Diffusion] Accelerate Cosmos3 Edge on Hopper with lossless fusions: #40386
- [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency: #40388
- [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (worker E2E 5.5 to 5.9% lower on one H200): #40405
- [Diffusion] Fuse lossless LingBot World FP32 normalization: #40425
- [Diffusion] Populate CPU weight stores before host registration: #40439
- [Diffusion] Add bounded exact conditioning cache across native models: #40470
- [Diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (worker E2E 7.4 to 8.8% lower on 2x H200 at TP2): #40486
- [Diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper: #40490
- [Diffusion] Fuse Joy Image Edit QKV concatenation and avoid QK copies: #40494
- [Diffusion] Support MiniMax-H3 PDD(Parallel Decoding Distillation) inference: #40568
- [Diffusion] Separate a use-scoped layerwise release from release_all: #40590
- [Diffusion] feat: allow a component use retain its layerwise resident set: #40592
- [Diffusion] Add a permanent lifetime for layerwise resident layers: #40599
- [Fix] Keep diffusion encoder TP context bindings consistent: #40646
- [Diffusion] Add opt-in SRT prompt enhancement to image and video APIs: #41095
- [Diffusion] fix: LoRA-wrapped linears crash Qwen-Image and MiniMax-H3 inference: #41272
- [Diffusion] Copy small files into overlay materialized trees instead of linking them: #41298
- [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200: #41305
- [Diffusion] Run an all-valid attention mask on the backend's unmasked kernel (Ideogram4 request latency 3.42 to 2.19 s at TP2 on B200): #41309
Local & Desktop AI
- [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts: #34528
- fix(moe): support Llama4 NVFP4 router input weights on SM120: #35504
- [PP][DeepSeek V4] Overlap communication and optimize SM120 prefill: #38792
- [Fix] Include SM121 in DeepGEMM packed-scale selection: #39482
- [Qwen4-Exp] Build the offloaded PLE table on the meta device so --ple-offload-embedding never materialises it on the accelerator: #39928
- Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090): #41292
AMD / ROCm
- [AMD] Speed up Wan2.2 DiT FP8 attention per-tensor quantization: #34695
- [AMD] Fix DSV4 FP4 dequant path for AITER on ROCm: #35123
- [ROCm][Diffusion] Enable fused qk norm and rope on ROCm: #35573
- [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4: #35619
- [AMD] Skip full-vocab softmax in EAGLE topk==1 draft on ROCm: #35872
- [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill: #36505
- MiniMax-M3: run the sparse prefill main attention through AITER Gluon paged attention: #36546
- MiniMax-M3: allocate the lightning-indexer K cache in fp8 on gfx95: #36549
- MoE: small-batch sorting path with fused mxfp8 quantisation: #36559
- MiniMax-M3: wave64 histogram-select decode top-k, and raise kMaxNumBlocks for CUDA graphs: #36560
- MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950: #36574
- MiniMax-M3: allow shared-experts fusion on ROCm gfx942 and newer: #36576
- [ROCm] Fix EAGLE spec-decode verify silently sampling greedy on HIP: #37134
- [ROCm] Widen the HiCache JIT copy rounds and enable the K-only host pool: #37152
- [AMD] Preserve deterministic inference when Lean Attention is enabled: #37740
- [AMD][Diffusion] FlyDSL fused norm kernels on wave32 targets (gfx1250): #37751
- [AMD] Fix DeepSeek-R1-MXFP4 accuracy with AITER FP8: #37762
- [AMD][DSV4] Enable hicache on deepseek-v4 fp8 unified attn: #37778
- [ROCm][DSV4] Enable breakable CUDA graph prefill: #37810
- [AMD] Enable GLM DSA prefill top-k to the v2 kernel (+4.9% token throughput at 70K input on GLM-5.2-MXFP4, 4x MI355X): #37889
- [AMD][Spec] Enable GDN ReplaySSM target-verify on ROCm: #38184
- [ROCm] Fuse the MLA q absorb into the RoPE + KV-write kernel on gfx950: #38340
- [AMD] Avoid the FP8 wo_a path when the weight is BF16: #38453
- [AMD][GLM-5.3-Flash Day 0] Route mHC through AITER on gfx950: #38545
- [AMD][GLM-5.3-Flash Day 0] Enable FP8 and Quark MXFP4 MoE on gfx950: #38546
- [AMD][GLM-5.3-Flash Day 0] Enable zero-RoPE TileLang DSA on gfx950: #38547
- [AMD][GLM-5.3-Flash Day 0] Honor fused and per-expert names in quark
exclude: #39317 - [AMD][GLM-5.3-Flash Day 0] Enable zero-RoPE MHA prefill on ROCm: #39338
- [AMD][GLM-5.3-Flash Day 0] Build the fused DSA k-pool top-k JIT kernel on HIP: #39339
- [AMD][GLM-5.3-Flash Day 0] Support non-2048 top-k widths in the DSA page-table transform: #39340
- [AMD][GLM-5.3-Flash Day 0] Enable the k-pool DSA indexer on gfx950: #39341
- [AMD][GLM-5.3-Flash Day 0] Enable speculative decoding (MTP) on ROCm: #39778
- [AMD][GLM-5.3-Flash Day 0] Load the MXFP4 MTP draft layer: #39779
- [AMD] Pad QSA MQA decode Q-heads to 16 for ROCm MFMA: #38875
- [AMD] Add a Triton packed sparse decode path for QSA on ROCm: #38876
- [AMD] Load fused shared experts for Qwen4-Exp and Qwen3.5 MTP: #38878
- [AMD][DSV4] feat: enable DSpark with fp8 unified_kv on gfx950: #38901
- [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe: #39059
- [ROCm][Fix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints: #39064
- [AMD][Kimi-K3] Fix deferred KDA gate projection and update DCP cookbook: #39066
- [AMD] GLM-5.2 NextN: cast draft fused MoE to per-channel FP8: #39155
- [AMD] Use exact CU share for gfx950 segment-plan headroom: #39503
- [AMD][Fix] Fix vattn_asm HIP error 709 under CUDA graph capture on ROCm 10: #39513
- [AMD] Fix deferred Kimi-K3 forget gate in fused in-projection: #39525
- [AMD][Fix] Fix DSV4 MTP crash: #39547
- [AMD] Prefer HIP Top-K for GLM-5.x on ROCm (P90 interactivity 96.97 to 106.69 tok/s/user on MI355X at concurrency 8): #39631
- [AMD] Update deepseek-v4 PDI and cache policy setting for agentic workload: #39702
- [AMD] Clamp MORI intranode grid GPUs: #39763
- [ROCm] fix: remove extra bf16 -> fp32 cast in jit grouped topk kernel path: #39775
- [ROCm] feat: enable aiter allreduce fusion for GLM models: #39790
- [AMD][Fix] Fix dsv4 server launch: #39875
- [AMD] Reuse KV gather indices across ASM context prefill layers: #39901
- [AMD] Pack Qwen3.5 GDN input projections on ROCm: #39902
- [AMD][DSV4] Allow moe_a2a_backend='mori' with DSpark + dp attention: #39910
- [AMD] Update ROCm AITER pin to acf8fdf9: #39965
- [AMD] dsv4: pick kv_splits per index stream, not by occupancy alone: #39968
- [AMD] Use Triton softmax routing for Qwen3.5 on gfx950: #39986
- [AMD] Tune Qwen3.5 TP4 GDN recurrent launch on gfx950: #39987
- [AMD][DSV4] fix: drop shadowing local get_exec import that breaks model startup on ROCm: #40186
- [AMD] Small-M MXFP4 fused-MoE kernel for gfx950 (Qwen): #40204
- [AMD] Drop the redundant scale zero-fill before AITER per-tensor FP8 quant: #40557
- [ROCm][DSA] Enable AITER fused FP8 indexer writer: #40710
- [AMD] Critical fix enabling Qwen3.8 FP8: restore dropped fused shared-expert weights: #40754
- [AMD] Add diffusion (Wan2.2) extras to gfx1151 Docker image: #40844
- [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens: #40878
- [AMD] Drop the unreachable vLLM fallback from ROCm FP8 activation quant: #40879
- [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe): #40943
- [AMD] Add tuned dsv4 shape: #40996
- [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt: #41120
- [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels: #41159
- [AMD][GLM5] Fuse shared expert into AITER MoE on gfx950: #41161
- [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm: #41377
NPU / Ascend
- feat(hicache): support Ascend Mamba states with FIA and async IO: #32500
- [NPU] decoding procedure optimization on qwen3.5/3.6: #35958
- [NPU] Fix xgrammar apply_vocab_mask device dispatch to use torch.ops.npu: #36120
- [NPU] add chunk gdn kernel and unify ssm state layout for ascend gdn backend: #36187
- fix(npu): fix hybrid KV transfer with PP prefill in PD disaggregation: #38402
- [NPU] Refactor weight processing and add NPUSwigluLimit activation: #38420
- [NPU][Fix] update low latency quantization input and update MXFP8 tests: #38831
- feat(npu): Support returning indexer top-k results: #39060
- [NPU][Diffusion] Optimize SenseNova-U1 batched generation: #39382
- [NPU] Avoid device synchronization in Ascend sampling: #39404
- [NPU] Support nccl backend for --remote-instance-weight-loader: #39413
- [NPU] Adapt hicache for K3 hybrid models: #39415
- dsv4(npu): support prefill context parallelism with interleave and zigzag: #39427
- [NPU][Diffusion] FA MXFP8 and modelslim w4a4f8 and w8a8f8 support for Wan2.2 and FLUX: #39438
- [NPU][Fix] Fix undefined require_attn_tp_allgather: #39561
- [NPU] support kimi k3 on A5 and improve performance: #39589
- [NPU] Support batch invariant FIA graphs for deterministic inference: #39607
- [NPU] fix npu docker workspace directory: #39732
- [NPU][DSV4] dsv4 enable cpp: #39820
- [NPU] Run arch35 block-FP8 dense linears on the native MXFP8 GEMM: #39823
- [NPU] Gate DFlash replay metadata refresh behind spec_algorithm check: #39879
- [NPU] Fuse MXFP4 W4A8 MoE gmm1 + swiglu + requant into one kernel: #39881
- [NPU] Avoid repeated BF16 wo_a weight transposes in DeepSeek-V4 decode: #39919
- [NPU][Fix] Avoid M-RoPE recompilation for variable sequence lengths: #40371
- [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU: #40446
- [NPU] Fix DSV4 hard-coding kv dtype: #40510
- [NPU] Update CANN version to 9.1.0: #40524
- [NPU] support NPU 910C L2 memcache offload: #41527
- [Docs][NPU] Add MiMo-V2.5-Pro FP4 DFlash best practice on Ascend NPU: #40577
CPU / Intel / XPU
- [Intel XPU] Enable fused_moe_triton tuning on XPU and add tuned DeepSeek-OCR-2 configs: #28723
- Speculative Decoding with NGRAM support for XPU: #31362
- [CPU] Add fp8_per_tensor_scaled_mm_cpu kernel: #32618
- [XPU] Enable HiSparse hierarchical sparse KV cache on Intel XPU: #32792
- [CPU][Diffusion] Add fused scale-shift and norm kernels for CPU: #33452
- [Intel][XPU][KVCanary] Enable KV Canary on Intel XPU: #33520
- [Intel][XPU] Enable chunked prefill scnearios for XPU with UT: #33804
- [XPU] Use torch scaled_mm for XPU block FP8 linear: #35605
- [Diffusion] Fix the XPU capability gates that broke the Wan2.2 A14B DiT path: #36825
- [XPU] Qwen3.8-flash-next enablement: #37213
- [CPU] Upgrade torch cpu dependencies to version 2.14.: #37939
- [XPU] weekly simple model enablement 2026/09/14: #39439
- [CPU] Add fused_sigmod_mul_cpu operators to the Meta Muse Glimmer model.: #39538
- [Intel GPU] Xpu/weekly simple model enablement 2026 09 21: #40664
- [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path: #40828
- [XPU] Bump sglang-kernel-xpu wheel to v0.3.0: #41220
Security
- fix(openai): reject request-supplied chat_template by default: #28135
- Restrict SafeUnpickler standard-library globals: #39858
- fix: restrict SafeUnpickler to explicit globals: #40259
Cookbook Updates
- DeepSeek-V4: a PD-disaggregated arm for B200 FP4 agentic serving with DSpark (#40610); MI355X Pro PD pairs with DSpark and UMBP, plus MegaMoE, FP8 KV attention, and breakable CUDA graph recipes: #41024, #41458
- GLM-5.2: MI355X MXFP4 recipes move to HIP Top-K and newer images, and the high-throughput cell adds HiCache: #40148, #40570, #41109, #41597
- GLM-5.3 and GLM-5.3-Flash: reasoning and tool-call parsers on by default, and the DCP option is temporarily removed from GLM-5.3-Flash: #40497, #40036
- Qwen3.5: GB200 and GB300 configs, and the MI355X HiCache recipe uses the kernel IO backend with the
page_firstlayout: #40770, #39572 - Qwen3.8-Flash-Next: NVIDIA NVFP4 on B200, B300, and GB300: #41046
- Kimi-K3: an Ascend 950PR/DT recipe: #40575
- Intel XPU: the cookbook now marks XPU-supported models: #33649
Dependencies
- xgrammar 0.2.1 to 0.2.7: #39026
- cache-dit 1.3.0 to 1.5.1: #40104
- sentencepiece pinned to 0.2.1 (0.2.2 rejects null pieces in InternVL tokenizers): #41777
- sgl-eval 0.1.0 to 0.1.2: #40620
- CUDA 13.4 image:
nvidia/cuda:13.4.1-cudnn-devel-ubuntu24.04base, sgl-kernel 0.4.7, sgl-deep-gemm 0.2.0, sgl-deep-ep 0.1.2: #40987; Triton 3.8 compatibility for DeepSeek-V4.1-Flash on Rubin: #40805 - CPU: torch 2.14.0, torchvision 0.29.0, triton 3.8.0: #37939
- XPU: sglang-kernel-xpu 0.2.0 to 0.3.0: #41220
- ROCm: AITER pin moves to
acf8fdf9: #39965 - NPU: CANN 9.1.0 with Python 3.12 (#40524) and sgl-kernel-npu 2026.9.0.post5: #40157
- Official releases now also publish a versioned renderer image: #40639
Breaking Changes & Upgrade Notes
- The unified radix tree uses the Rust core by default. Configurations the Rust core does not support fall back to Python automatically; set
SGLANG_UNIFIED_RADIX_TREE_CORE_BACKEND=pythonto force the Python core everywhere: #39627 - FlashInfer MoE fused finalize is off by default for numerical accuracy. This costs some W4A4 MoE latency; set
SGLANG_FLASHINFER_MOE_FUSED_FINALIZE=1to restore the fused atomic finalize: #40105 - PD decode now defers KV release on abort by default on the Mooncake and NIXL backends, holding pages until every prefill rank acks (30 s timeout,
SGLANG_DISAGGREGATION_DEFERRED_DECODE_KV_RELEASE_TIMEOUT). SetSGLANG_DISAGGREGATION_DEFERRED_DECODE_KV_RELEASE=Falseto restore immediate release: #41023 - Request-supplied chat templates are rejected. A
chat_templatekey insidechat_template_kwargson/v1/chat/completionsor/v1/tokenizenow returns an error unless the server starts with--trust-request-chat-template; otherchat_template_kwargsare unchanged: #28135 - Deprecated endpoints, env vars, and aliases are removed. Use
/model_infoinstead of/get_weight_versionand/weight_version,/v1/loadsinstead of/get_load, andPOST /hicache/storage-backend/clearinstead of/clear_hicache_storage_backend.SGLANG_NSA_*aliases give way toSGLANG_DSA_*,SGLANG_ENABLE_SPEC_V2is gone,--elastic-ep-rejoinbecomes--elastic-ep-join-mode recover,--kv-cache-dtype fp4_e2m1becomesfp4_mx_block16, andpython -m sglang.bench_offline_throughput/sglang.bench_one_batchmove tosglang.benchmark.offline_throughput/sglang.benchmark.one_batch: #40795 - Legacy radix cache implementations are removed. The experimental C++ radix tree (
SGLANG_EXPERIMENTAL_CPP_RADIX_TREE) is gone (#40775), andSGLANG_ENABLE_UNIFIED_RADIX_TREEno longer does anything, though it still prints a deprecation warning (#40780). Code importingSWARadixCache,MambaRadixCache, orHiRadixCachemust move toUnifiedRadixCache: #40313, #40787 - Custom logit processors must return full-shape logits. A
CustomLogitProcessorthat returned a reduced shape such as(1, vocab)and relied on broadcasting now raises at decode time. The built-in processors are unaffected: #39234 - RL weight updates run inside a session.
update_weights_from_tensorandupdate_weights_from_distributedmust be called betweenbegin_weight_update()andend_weight_update().UpdateWeightsFromTensorReqInput.disable_draft_modelis replaced byselector="target", andChecksumInfo.parallelism_infois now a list with one entry per role: #40777 - Parallel-state getters are deprecated. They are no longer re-exported from
sglang.srt.distributed; read groups throughget_parallel()(for exampleget_parallel().tp_group) or import fromparallel_statedirectly: #40342, #40344, #40707 - Ascend A2 (910b) images are no longer published. Release and nightly images ship for A3 and 950 (A5) only, on CANN 9.1.0 with Python 3.12: #40524