github sgl-project/sglang v0.5.21

2 hours ago

Highlights

779 PRs from 227 contributors.

New models

Model Type Cookbook
DeepSeek-V4.1 Flash LLM / VLM link
GigaChat 3.5 LLM / VLM link
IQuest-Q1 LLM / VLM link
MiMo-V2.6 / MiMo-V2.6-Pro LLM / VLM link
Ling-3.0-flash-VL LLM / VLM link
DiffusionGemma Diffusion link
Qwen-Image 2.1 Diffusion link
Anima Base v1.0 Diffusion link
Ming-Image 0.1 Design / Design-Layer Diffusion link
FLUX 3 Action Diffusion link

Key features

  • PD instances can switch between prefill and decode on the fly, no restart needed (#28403).
  • The prefix cache now runs on a Rust core by default (#39627).
  • DeepSeek-V4.1 gets 22% faster first token on long prompts (#40352).
  • Kimi K3 gets 20.6% higher prefill throughput in PD serving (#40045).
  • More accurate results under PP, DP attention, and CP, with layer communication now handled by SGLang (guide, #41557).
  • New Decisions API (/v1/decisions) turns an LLM / VLM into a low-latency classifier and scorer (docs, #41208).
  • Score API (/v1/score) scores all candidates in one request (#38965, #41188).
  • Run MiniMax-H3 with SGLang Diffusion inside ComfyUI (#35990).

To upgrade

uv pip install --prerelease=allow sglang==0.5.21
Platform Docker image
NVIDIA (CUDA 13) lmsysorg/sglang:v0.5.21
AMD MI35x lmsysorg/sglang:v0.5.21-rocm10-mi35x
AMD MI30x lmsysorg/sglang:v0.5.21-rocm10-mi30x
Intel GPU lmsysorg/sglang:v0.5.21-xpu
Intel CPU lmsysorg/sglang:v0.5.21-xeon

Full Release Notes

Speculative Decoding

  • Pipeline parallelism x speculative decoding (EAGLE/MTP) compatibility: #30775
  • Support XQA backend for SpecDec verify: #32269
  • [Spec] Windowed draft-decode attention for built-in EAGLE / MTP drafts: #32673
  • Avoid materializing GDN QKV tensors during target verification: #33778
  • [Spec] Fix CDF boundary handling in TreeSpeculativeSamplingTargetOnly: #35798
  • [Spec] Add LiLiCorr: a candidate-lattice reranker for DFlash drafts: #37462
  • Add out-of-tree DFlash extension points: #38740
  • [Spec] Add explicit prefill shared-read capability for plugins: #39502
  • [Fix] Don't write conv state from the fused KDA verify kernel: #39524
  • [Spec] Model-agnostic last-stage draft embedding under pipeline parallelism: #39643
  • Use runtime token widths for Triton speculative verification: #39859
  • avoid host sync in DSpark prefill slot expansion: #40111
  • [Experimental] Preserve speculative decoding during prefill across DP ranks: #40118
  • [Spec][PP] Launch extend microbatches before the spec output exchange: #40499
  • [KDA] Enable ReplaySSM for GLM-5.3 Flash: #40517
  • [LFM2-VL] Add DSpark speculative decoding (1.66x to 2.56x speedup at batch 1): #40651
  • [Spec] Support DFLASH for Kimi K3: #40794
  • [Fix] Reduce Nemotron MTP attention outputs once: #40800
  • [Fix] Capture complete Nemotron auxiliary hidden states: #40801
  • [Fix] Skip the DCP target-verify MLA kernel during FlashInfer autotune: #41138
  • Fix mixed chunk prefill with DP speculative coordination: #41179
  • [Fix] Plan NextN / MTP draft layers as one-layer models and fix the Bailing V2 NextN draft: #41194

Piecewise & Breakable CUDA Graph

  • [Fix] Preserve model runner contracts in prefill CUDA graphs: #35452
  • [DSV4][BCG] Optimize the heavy memory use of C4 Indexer when BCG is enabled (capture memory down 58% to 18 GB): #36534
  • [Fix] Don't free the multi-CTAs KV counter the decode graphs captured: #39175
  • [CPU] Avoid prefill CP predicates during decode graph capture: #39690
  • [Runtime] Add decode CUDA graph hooks for eager logits processing: #40222
  • [Kimi K3] Fix CUDA graph stream explosion: #40640
  • [DSpark] Fix draft CUDA graph stream explosion: #40658
  • [Fix] Give the full prefill CUDA graph replay view the captured bucket's input_ids: #40851
  • [DSA] Fix the pooled-indexer breakable prefill bridge under DP attention: #41311

Attention Backends

  • Allow attention layers to opt out of the prefill wrapper: #40683
  • [Fix] Use cached prefix lengths for FlashInfer full-attention ragged prefill: #40796

MoE & Expert Parallelism

  • [3/N] elastic-ep: Recapture decode CUDA graphs after scale-up: #33723
  • Fix int4 MoE tuner config filename: #35260
  • [MegaMoE] Wire Qwen MoE blocks to DeepGEMM MegaMoE (MXFP4 and NVFP4 experts): #38080
  • [Feature] Support BF16 and batch-invariant inference with DeepEP v2: #38160
  • [Perf] Optimize w4a8 MoE for glm5.2 on H200: #38220
  • [Fix] Guard FlashInfer CUTLASS MoE against 0-token inputs: #38780
  • [Kernel] Add H20 block-FP8 MoE configs for GLM-5.3-Flash EP4/EP8: #38913
  • Accept MXFP8 dispatch in FlashInfer A2A TRT-LLM MoE: #39613
  • [MoE] Honor swiglu_limit clamped activation in flashinfer_cutlass runner: #39939
  • Support MXFP8 and deferred route weighting in DeepEP v2: #40030
  • [MoE] Disable FlashInfer fused finalize by default for numerical accuracy: #40105
  • [Fix] Fix top-1 MoE routing with non-unit scaling: #40187
  • [DeepEP v2] support GLM-5.3-Flash (Glm5NextForConditionalGeneration): #40466
  • [Fix] Decide the MoE padded-row bound from the layer scatter mode: #40672
  • [DeepEP v2] Let a model package supply its per-rank prefill dispatch bound: #41201

Quantization

  • [MXFP8-KV] Skip writes to the reserved CUDA-graph padding slot: #35351
  • feat(kv-cache): support SM100 NVFP4 GenMHA and speculative decoding (1.355x faster decode attention kernel at 1M context): #36340
  • [Quant] ModelOpt mixed precision: dispatch block-FP8 MoE experts and derive the block size: #38726
  • fix(modelopt): dispatch NVFP4 MoE on the cached backend, not the live global: #38932
  • [Quant] Serve 32-wide-K ue8m0 block-FP8 linears through the FlashInfer MXFP8 GEMMs: #40039
  • [ModelOpt][PP] Keep BF16 shared experts out of the NVFP4 fusion so TP1 pipeline stages can load: #40628

Parallelism & Disaggregation

  • [PD] Introduce runtime role switching between prefill and decode: #28403
  • disaggregation: Fix FakeKVSender queue accumulation: #28652
  • [PD] Keep EAGLE DP graph and token metadata consistent: #32196
  • [NVIDIA][comm] Merge EP+MoE-TP post-experts all-reduces into one _TP reduction: #32963
  • Add 8-node AllReduce/AllGather and MNVLS algorithm support to MSCCL++: #37442
  • [KV-Shard 2/4] Sharded pools: #37615
  • [KV-Shard 3/4] Enable Control plane: #38468
  • [KV-Shard 3/4] Enable Control Plane B: #39964
  • Fix device context during NIXL backend initialization: #38774
  • [PD] Do not admit intake-rejected requests to a PD handoff: #38935
  • Reduce decode bootstrap latency with request-owned speculative KV (mean TTFT 19.4 to 15.1 s on DeepSeek-V4-Pro 1P1D, MI355X, concurrency 256): #38978
  • Fix disagg PP MTP for GLM-5.2: #39378
  • [PD] Add optional KV transfer checksums: #39500
  • [DCP] Use logical token capacity for PD admission and load reporting: #39731
  • Refactor the Cute-DSL AR fusion to support DeepseekV2 archs (GLM-5.3, etc.): #39816
  • [PD] Validate Mooncake EFA allocator compatibility: #39973
  • [PD] Skip singleton transfer-status all-reduces: #40003
  • [PD] Enable optimistic prefill with buffer-only L3 write-through HiCache: #40043
  • [Kimi K3][PD] support pp prefill + dcp decode with dspark: #40045
  • [PD] Add decode host receive for custom transfer backends: #40238
  • [PD] Allow decode radix cache and HiCache L1/L2 with DCP: #40263
  • [PD] Enter the custom mem pool once when allocating DCP pack buffers: #40284
  • [PD] Bound cached-prefix DCP transfers by pack capacity (cached-prefix throughput about 10.9x higher at concurrency 8 on Kimi Linear, 8x B200): #40376
  • [PD] Pack draft KV head slices for DCP transfers: #40500
  • [PD] Preserve abort ACKs until in-flight KV transfers drain: #40645
  • Fix NIXL transfer of MXFP8 KV block scales: #40792
  • [PD] Honor gracefully_exit in disaggregation event loops and keep non-zero-rank launchers alive on SIGTERM: #40793
  • [DP Attention] Publish DP buffer sizes from a ForwardBatch: #40858
  • [PD] Enable deferred decode-side KV release by default: #41023
  • [PD] Add a none decode retraction backup and subclass seams in the PD queues: #41103
  • [PD] fix: cache resumed decode-radix requests from root instead of an unlocked re-match: #41261
  • [PD] Fan drain abort ACKs out to every decode peer of the room: #41402
  • [PD] Defer decode KV release on every transfer failure, not only decode-initiated aborts: #41404
  • [Fix] Stop deferring the last layer's FFN all-reduce in five models: #40868
  • [Fix] Complete the deferred FFN all-reduce before a pipeline-parallel send: #41079
  • [Fix] Complete the deferred FFN all-reduce before deepstack addition and aux hidden-state capture: #41080
  • [Fix] Step-3.5: stop dense layers from summing their output twice under DP attention: #41082
  • [Fix] Broadcast requests along attention CP before attention TP: #41083
  • [Fix] Complete the all-reduce when the flashinfer fused norm declines a batch: #41193
  • [Fix] Stop counting a deferred FFN sum more than once: replicated TP1 shared expert, dense reduce_scatterv: #41195
  • [Fix] Keep one copy of CP-replicated rows in the DP gather: #41422
  • [Fix] Gather a dense FFN's input across attention DP and CP in one DP sum: #41423
  • [Fix] Run MoE layers under attention DP and GQA prefill CP on the declared DP × CP gather: #41432
  • [Fix] Falcon-H1: count the Mamba mixer's output once under tensor parallelism: #41433
  • [Fix] LongCat-Flash under attention DP: branch and merge the dense FFNs through the communicators: #41436
  • [Refactor] Group communicator fusion and CP adapters: #41547
  • [Refactor] Centralize decoder output access: #41548
  • [Refactor] Carry residual state across stage boundaries: #41549
  • [Refactor] Capture auxiliary states at residual reads: #41550
  • [Refactor] Select reduction fusion at the consumer: #41551
  • [Refactor] Construct independent decoder stage boundaries: #41552
  • [Refactor] Migrate specialized decoder and overlap boundaries: #41553
  • [Refactor] Retire the layer facade and simplify boundary internals: #41554
  • [Refactor] Rename the module to layer_boundary: #41555
  • [Refactor] Group layer boundary unit tests: #41556
  • [Refactor] Document layer boundary contracts and integration: #41557

Scheduler & Runtime

  • feat: support custom OTLP trace service name: #35802
  • [RL] Release the weight-checker snapshot once compare passes: #37284
  • [Runtime] Let out-of-tree platforms provide full graph backends: #37969
  • Fix: post-load staging regression breaks offload meta/sharded_gpu modes: #38779
  • feat: add kv hint envelope to request transport: #38891
  • [Observability] Fix negative queue_time for retracted requests: #39312
  • Fix first-token metadata and reused attention-layer indexing: #39328
  • perf(engine): avoid timed waits for Engine responses (+37.7% output throughput in a Dynamo TCP push benchmark at the default batch_notify_size): #39486
  • [Misc] Fix tool-call index, graph padded-row count, and prefill-graph input_embeds refresh: #39574
  • [Misc] Merge FlashInfer autotune caches across spec workers, pad MXFP4 TP shards, drop dead ngram attrs: #39678
  • [Observability] Expose python/rust frontend identity in /server_info: #39993
  • [Metrics] Propagate idle gaps across all scheduler loops: #40004
  • [Scheduler] Count complete prefill bursts and their tokens: #40006
  • [Scheduler] Add shortest-prefill-first scheduling: #40024
  • [Profiler] Ignore PREBUILT batches in profile-by-stage: #40098
  • Enable optimistic prefill for Mamba radix-cache models: #40184
  • [Perf] Fork-safe import: no CUDA context at import time, lighter argument parsing: #40201
  • [Feature] Support --tokenizer-worker-num > 1 in the offline Engine API: #40260
  • Clean up startup logging and streamline log audits: #40526
  • [Fix] Run KV canary hooks for context-parallel prefill: #40642
  • [RL] Fix Kimi K3 expert-count lookup for routed-expert capture: #40700
  • [RL] Add RL weight-update sessions and support updating spec draft runners: #40777
  • [RL] Keep pause_generation and weight updates from deadlocking each other: #40779
  • [Misc] Remove deprecated endpoints, env vars and aliases past two releases: #40795
  • [Metrics] Log forward and forward+idle occupancy over total wall time: #40802
  • [RL] Keep DSA cuda-graph state and the graph pool intact across TMS pause/resume: #40804
  • [Fix] Recover from stale torch extension locks in every cpp_extension loader: #40989
  • [Perf] Lazy-load built-in model definitions and nixl_ep at startup: #41061
  • [Model Loader] Stop checkpoint prefetch after iterator completion: #41588
  • [Benchmark] Limit warmup concurrency in serving benchmark: #39398
  • Support LongBench v2 in one-batch server benchmarks: #39874
  • [Benchmark] Take each request's prompt length from the server: #39889
  • [Benchmark] Add agentic rollout simulator and offline explorer: #40034
  • [Benchmark] Optionally clear HiCache storage between cases: #40659

Sampling

  • fix(sampling): validate sampling_seed is an int within int64 range: #28960
  • perf(sampling): avoid GPU syncs when applying custom logit processors: #39234
  • Use pinned memory for asynchronous sampling metadata transfers: #39777
  • [Logprob] Borrow graph-pool memory for input logprob logits construction: #40007
  • [Logprob] Serve input-logprob temporaries from CUDA-graph-pool dead space: #40038
  • [Sampling] Add selected/support sampling logprob modes: #40932
  • [Sampling] Stream sampling masks as per-request arrays: #40986
  • [PD] Keep the sampling mask of a replayed rebootstrap token: #41235

HiCache & Radix Cache

  • [Radix Cache] Sync Rust TreeCore and make it the default: #39627
  • Add Agentic-Aware Tail-Optimized LRU eviction to the unified radix cache (opt-in; p99 ITL down 43.9% at concurrency 32 on DeepSeek-V4-Pro FP4, 4x B300, agentic trace): #34012
  • Feat: Add TensorCast storage as a new HiCache backend: #27265
  • [HiCache] fix: resolve Mooncake local_hostname per node for runtime attach: #29668
  • [Fix] fix(hicache): wait for decode offload before retraction: #30899
  • [HiSparse] Add MHA hisparse support for MiniMax M3: #31446
  • [Mooncake] Fix silent SSD offload corruption when TP/PP ranks share ssd_offload_path: #31926
  • [Scheduler] Align RadixCache no-insert cleanup with kv_len_to_handle: #35204
  • [PP + HiCache] Add PP Prefetch Tickets for eager cross-stage storage prefetch: #36700
  • Use a shared byte budget for unified hybrid-SWA memory: #36729
  • Gate mamba extra-buffer predicates on uses_mamba_radix_cache: #37474
  • [Unified Memory] Hierarchical cache for every unified pool shape: #37507
  • [HiCache] Fix sparse hybrid transfer layer IDs: #37870
  • [Unified Memory] fix: preserve FP8 dtype in unified MHA pool: #38133
  • [HiCache] Yield idle scheduler so storage workers can drain: #38504
  • [LMCache] Support lmcache unified radix cache: #38652
  • [Unified Cache] Dedup replicated MLA/DSA KV in the UMBP direct linker: #38778
  • [Perf] Fuse SWA page lookup and mapping clear: #38948
  • [HiCache][Perf] fix: batch HiCache D2H submits per step for hybrid pools: #39050
  • [HiCache] Back up MXFP8 KV scales in the host pool: #39089
  • [Mamba] Fix checkpoint depth for prefixes that end off the radix page: #39115
  • [HiCache] Label radix-cache metrics per rank and split the "shrunk" prefetch reason: #39280
  • bugfix:fix unifiedcache c128 radix cache management: #39426
  • Support unified memory page-envelope transfers in PD: #39477
  • Support unified memory decode host pools: #39478
  • [Unified Cache][9/N] add opt-in MLA load deduplication for Mooncake Linker (TTFT 3.85 to 2.23 s with a 256K cached prefix, GLM-5.2 on 8x H20): #39565
  • [HiCache] Forward prefix metadata to v2 storage calls: #39567
  • [HiCache] Keep hybrid transfer layer maps stage-local under PP: #39699
  • [HiCache] Add the page-unified KV load-back JIT kernel: #39726
  • [2/N][Kernel] Fuse padding-preserving HiSparse slot translation: #39837
  • [Unified Tree] fix: exempt host-locked aux nodes from the sanity_check host-LRU check: #39980
  • [HiCache] Read the in-flight buffer backup's node id from its snapshot in sanity_check: #40013
  • [HiCache] Stop arming a prefetch retry for a too-short storage span: #40042
  • [MemCache] Release up to owned_kv_len on radix cache insert: #40075
  • [HiCache] Auto-size the host pool to fit available host memory: #40135
  • Preallocate HiCache MHA staging before post-capture KV sizing: #40256
  • Fix prefetch attempt cleanup on abort: #40262
  • [HiCache] TMA-staged host<->device KV transfer kernel (sm_90+): #40278
  • [HiCache] Size MHA host pools from device row width: #40304
  • [Fix] Missing SWA eviction during decode preallocation: #40309
  • [HiCache] fix: bound the controller reset join so a stalled storage thread cannot hang the scheduler: #40312
  • Remove swa and mamba radix cache: #40313
  • [Fix] Forward SWA prealloc reclaim through the DSV4 HiSparse allocator: #40354
  • [HiCache] Give trailing sidecar storage transfers a contiguous prefix_keys chain: #40456
  • [HiCache] Make host reclamation independent of transfer order: #40512
  • [HiCache] Demote internal-node mamba states on write_back eviction: #40680
  • [HiCache] Demote internal-node SWA KV to host on write_back eviction instead of dropping it: #40712
  • [MemCache] Remove the experimental C++ radix tree: #40775
  • [MemCache] Clean up SWA/Mamba radix cache leftovers and drop SGLANG_ENABLE_UNIFIED_RADIX_TREE: #40780
  • [HiCache] Remove the unused HiRadixCache: #40787
  • [MemCache] Free the rows below the SWA evict floor on all-SWA request release: #40798
  • [HiCache] Batch buffer-only KV backups within each flush: #40960
  • [Fix] Derive per-runner hybrid SWA layer ids on ModelLayerInfo instead of mutating ModelConfig: #40983
  • [HiCache] fix: Drain pending backups before internal Mamba write-back: #41092
  • [Unified Memory] Fix Inkling conv-checkpoint track ids written to virtual slot numbers: #41144
  • Let predicate-registered linear-attention models carry the mamba radix-cache leaves: #41165
  • [Unified Memory] Honor move gates in float relocation and size auto HiCache from host capacity: #41248
  • [MemCache] Unify component eviction cursors and lock receipts: #41276
  • [MemCache] Never free the protected prefix on request release: #41312
  • Make sliding-window caching and speculative batch padding extensible: #41325
  • [MemCache] Fix LMCache component cursors and per-cache backend selection: #41328

LoRA

  • [LoRA] Support DP attention in LoRA backends: #36389
  • [LoRA] Size dense row/column-parallel LoRA buffers from the base linear's real shard: #39379

Multimodal

  • fix(multimodal): return 400 for corrupt image inputs: #28131
  • [Multimodal] Avoid CUDA placement on non-CUDA platforms: #38750
  • Fix Mistral3 retaining every vision-tower layer to read one: #39185
  • perf(multimodal): offload CPU feature hashing with bounded admission: #39539
  • Carry deferred attention operands and reuse multimodal shared memory: #39870
  • [MM] Skip VMM error gathers for text-only requests: #40005
  • [MM] Copy placeholder ids to CUDA asynchronously: #40010
  • [MM] Keep scheduler padding in packed token arrays: #40357
  • [Fix] Raise on undelivered embeddings in send_with_url, fix broken tests: #40502
  • Fix multimodal feature offload races: #40621
  • [VLM] Introduce FA4 into ViT for SM100/SM103: #41344

Model Support & Optimizations

  • [Feature] Gigachat 3.5 support: #29189
  • [dLLM] Support DiffusionGemma serving: #34061
  • Add Ling-3.0-flash-VL model support: #38526
  • dsv4.1: remaining model and runtime integration: #38798
  • [DSv4.1] Score prefill consumer index layers on candidate blocks with DeepGEMM: #40352
  • [Feature] Xiaomi MiMo-V2.6/MiMo-V2.6-Pro day0 support: #40448
  • [Model] Add IQuest Q1 support and MTP draft: #41590
  • fix(gemma4): set lm_head_is_tied for Gemma4UnifiedForConditionalGeneration: #35809
  • [Fix] Preserve YaRN scaling when extending rotary caches: #38786
  • [Kimi-K3] O(1) expert weight lookup in load_weights (expert-name matching 21.45 to 0.13 s at load): #38805
  • [Model] Serve DeepSeek-OCR-2 with its official 768px local-crop geometry: #38996
  • Fix GLM-OCR MTP multimodal embeddings and positions: #39088
  • [DSV4] Chunk the indexer MQA logits by query rows under a free-memory budget: #39095
  • [Perf] Fuse the glm5_next mHC attn->MLP boundary: #39200
  • [DSV4] fix: keep the TileLang JIT cache under SGLANG_CACHE_DIR: #39364
  • Fuse GLM-5.3-Flash KDA projections and prefill metadata: #39688
  • [GLM-5.3-Flash] Reduce KPool planning synchronization and overlap indexer preparation: #39695
  • [Fix] Fix GLM5 mHC PP forward: #39720
  • [DSV4] Generalize attention metadata, sparse prefill, and KV pool over compress ratios: #39921
  • [Qwen3.8 Next] Fuse Qwen PLE gate and convolution preparation for target verify: #40041
  • Support GLM-5.3-Flash hybrid attention CPU offload and PD index mapping: #40310
  • [DSV4] fix: size the C4 state ring by the page it is addressed by: #40337
  • [Fix] Keep mHC context out of non-V4 compiled MoE forwards: #40353
  • [Qwen3.8-Next] Pipeline-parallel serving and PD-prefill MTP for Qwen4-Exp: #40501
  • Fix GLM-5.3 forget-gate shape for nvCUTEDSL verify: #40607
  • [Fix] Avoid duplicate residual in LongCat MoE shortcut: #40799
  • [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget: #40854
  • [DSV4] Budget the ratio-2 pair state pool in DSV4PoolConfigurator: #41048
  • [DSV4] Size compressed pools from one per-ratio table in DSV4PoolConfigurator: #41049
  • [Fix] Add name mapping in load_weights of nvidia/LocateAnything-3B: #41059
  • [DSV4] Fix TRTLLM uniform FP8 KV memory budgeting: #41090
  • [DSV4] Account for FlashMLA physical KV page padding in memory budgets: #41091
  • [Kimi-K3] Merge fused_qkvg_proj into the loader-seeded packed_modules_mapping: #41164
  • [Qwen3.8 Next] Fuse small CUDA graph input buffer copies: #41166

Kernel Library

  • [Kernel] Share the warp vectorized copy and enforce its alignment: #36176
  • fix(kda_prefill): fence shared writes before async proxy reads: #39124
  • [Kernel] Coalesce the KDA CuTe DSL decode state transpose: ~3x faster, bit-identical: #39680
  • [Kernel] Move CUDA and ROCm speculative kernels to JIT: #40033
  • Fix TopK v2 fallback when 16-block cluster capacity is zero: #40163
  • [Kernel] Fuse hc_combine_norm for mid-size verify batches (9-96 rows): #40208
  • [KDA] Fix missing beta sigmoid in PTX prefill: #40685
  • [JIT] Add an occupancy-preserving L1 carveout preference: #40767
  • [Perf] Mamba2 selective_state_update up to 2x faster on B200 via 8x1 launch config for dstate 128 (+7.5% Nemotron-3-Super serving): #41223

Configuration System

  • Add registration for external model configurations: #39452

OpenAI-Compatible API

  • [Fix] Guard conditional top-logprob keys in the completions echo path: #34776
  • [Fix] Prevalidate JSON Schema support per grammar backend: #37839
  • fix(openai): recover logprobs token bytes from token_id (UTF-8 fragments): #38604
  • [Score API] Setwise Scoring Support: #38965
  • feat: use XGrammar V4.1 DSML parameter constraints: #39026
  • fix(function_call): buffer complete DeepSeek DSML invokes: #39632
  • Fix corrupted chat prompts on mistral_common tokenizers (tool_choice auto never fires): #39773
  • [Fix] Keep Inkling automatic tool grammar active across the response: #40468
  • fix(grpc): expose native response timeout as a server argument: #40644
  • [Feature] Add per-item candidate token scoring and calibration: #40826
  • Add MiniMax arch fallback to auto parser resolution: #40930
  • [Score API] Setwise scoring: CausalLM support (batched + --enable-mis): #41188
  • [Feature] Add a System One compatible /v1/systemone route (also ships the /v1/decisions route it builds on): #41208
  • Fix chat template cache key order: #41517

Rust Server

The experimental Rust router and frontend (experimental/sgl-router, the Rust renderer) keep moving toward parity with the Python server.

  • [Rust Renderer] Standalone preprocessing: #36718
  • [SGL Router] Render chat prompts with dynamo-render: #38983
  • [Router] Fleet-wide sampling contract 1/3: the config surface: #39000
  • [Router] Fleet-wide sampling contract 2/3: enforce and inject per request: #39001
  • [Router] Fleet-wide sampling contract 3/3: splice injection without re-serializing: #39002
  • [Router] Resolve a wire protocol per worker at registration: #39004
  • [Router] Speak cleartext h2c on both edges: serve it inbound, forward it outbound: #39006
  • [Router] Count open HTTP exchanges until their response body finishes: #39014
  • [Router] Add the shutdown-drain configuration surface: #39015
  • [Router] Drain readiness before SIGTERM shutdown so k8s deregisters the pod first: #39016
  • [Router] Real-GPU e2e coverage for storage-tier-aware cache routing (3/4): #39110
  • [Router] Add a KV storage-tier coverage row to the Grafana dashboard (4/4): #39111
  • [Router] Improve SGLang chat render parity: #39133
  • [Router] Shard the cache-aware KV tree by chain root: #39167
  • [Router] Add --worker-queue-limit: stop sending cache-affinity traffic to a queueing worker: #39168
  • [Router] Pin to the prefix owner when the whole fleet is queueing (--saturation-queue-floor): #39169
  • [Router] Sample k random candidates for the min-load fallback (--min-load-choices): #39170
  • Add external multimodal processors to the Rust frontend: #39329
  • [Rust] Extract a transport-neutral frontend core: #39385
  • [SGL Router] Forward input_ids only for string content; count tokenize errors only when forwardable: #39458
  • [Router] Abort the engine when a client disconnects mid-request: #39461
  • [Router] Derive error status from a failure class; preserve the worker's status (1/3): #39463
  • [Router] Treat an upstream 503/429 as backpressure, not a breaker fault (2/3): #39464
  • [Router] Log every request at one site; derive its outcome from the final status (3/3): #39465
  • [SGL Router] Share model-file discovery for chat formatters: #39485
  • Rust server unify datapath for mm and generate requests: #39679
  • [SGL Router] Add Kimi-K3 rendering with SGLang parity: #40390
  • [SGL Router] Bound streaming lifetimes and release guards on idle disconnect: #40391
  • Port chat_parsing core: #40477
  • [SGL Router] Match DeepSeek V4 rendering to SGLang: #40530
  • [SGL Router] Add SGLang-compatible DeepSeek V4.1 Flash rendering: #40532
  • [SGL Router] Release cancelled circuit-breaker probes: #40603
  • [SGL Router] Fix readiness, IPv6 discovery, logging, and model validation: #40604
  • [Router] Give the cache-aware tree a snapshot surface (1/13): #40687
  • [Router] Serve the cache-aware tree at /internal/kv_snapshot (2/13): #40688
  • [Router] Name a replica's siblings with --kv-peer-selector (3/13): #40689
  • [Rust Renderer] decouple renderer sampling from protocols: #40747
  • [SGL Router] Launch reorg routing with existing policy options: #40766
  • [SGL Router] Track input_ids forwarding outcomes per chat request: #41185
  • [SGL Router] Take SGLang's render defaults: --default-chat-template-kwargs and thinking/effort envs: #41221
  • [SGL Router] Scope input_ids forwarding by renderer: all text chats for DeepSeek-V4: #41226
  • [Rust Frontend] Decode input_ids without untagged buffering (input_ids deserialization about 45% faster): #41246
  • [SGL Router] Book the input_ids forwarding outcome only for built bodies: #41342

Simulator

  • [Simulator][Compatibility] Adapt to latest KV cache pool interfaces: #40418

SGLang-Diffusion

  • [Diffusion] model: support qwen-image-2.1: #39983
  • [Diffusion] Add native Anima Base v1.0 support: #41011
  • [Diffusion] Support FLUX 3 Action robot policies: #41066
  • [Diffusion] Add native Ming-Image Design and Design-Layer support: #41067
  • [Diffusion] add /metrics support: #19084
  • [Diffusion] Preserve per-sample rollout trajectories across multi-output merge: #34416
  • [Diffusion] Fix AttributeError in grouped forward_batch by installing the residency manager: #34417
  • [Diffusion] Apply latent-ids and packing to caller-provided initial latents: #34418
  • [Diffusion] MiniMax-H3 Spectrum skip-step + fused RMSNorm/AdaLN: #35684
  • [Diffusion] Add MiniMax-H3 to ComfyUI integrated mode: #35990
  • [Diffusion] In-place LoRA merge/unmerge under layerwise offload: #36192
  • [Diffusion] feature: out of tree platform support: #37547
  • [Diffusion] support multiple task types for pipelines: #38762
  • [Diffusion] Enable shared RMSNorm dispatch for SenseNova-U1: #39705
  • [Diffusion] Preserve explicit attention backends during autotune: #39882
  • [Diffusion] Cache-DiT 1.5.1: DMD Calibrator, SVDQuant DQ, etc.: #40104
  • [Diffusion][MiniMax-H3] Add SM120 Sage compute for SubBlock sparse attention: #40116
  • [Diffusion] attention: add fp8_fa_sm120 FP8 backend for SM120 GPUs: #40175
  • [Diffusion] Fuse lossless SenseNova RoPE for 5% faster H200 inference: #40374
  • [Diffusion] Fuse rounded SwiGLU for quantized MiniMax-H3 MLPs (FastH3 FP8 worker E2E 1.9% lower on 2x H200): #40378
  • [Diffusion] Fuse LongCat GELU+cat and support Edit-Turbo BCG: #40384
  • [Diffusion] Accelerate Cosmos3 Edge on Hopper with lossless fusions: #40386
  • [Diffusion] Enable lossless SANA-Video eager conv fusions for 12.6% lower latency: #40388
  • [Diffusion] Fuse lossless Wan VAE post-ops for LongLive 2 I2V (worker E2E 5.5 to 5.9% lower on one H200): #40405
  • [Diffusion] Fuse lossless LingBot World FP32 normalization: #40425
  • [Diffusion] Populate CPU weight stores before host registration: #40439
  • [Diffusion] Add bounded exact conditioning cache across native models: #40470
  • [Diffusion] Enable lossless Cosmos3 Super T2I QK fusion on Hopper TP2 (worker E2E 7.4 to 8.8% lower on 2x H200 at TP2): #40486
  • [Diffusion] Fuse lossless Klein packed QK RMSNorm and RoPE on Hopper: #40490
  • [Diffusion] Fuse Joy Image Edit QKV concatenation and avoid QK copies: #40494
  • [Diffusion] Support MiniMax-H3 PDD(Parallel Decoding Distillation) inference: #40568
  • [Diffusion] Separate a use-scoped layerwise release from release_all: #40590
  • [Diffusion] feat: allow a component use retain its layerwise resident set: #40592
  • [Diffusion] Add a permanent lifetime for layerwise resident layers: #40599
  • [Fix] Keep diffusion encoder TP context bindings consistent: #40646
  • [Diffusion] Add opt-in SRT prompt enhancement to image and video APIs: #41095
  • [Diffusion] fix: LoRA-wrapped linears crash Qwen-Image and MiniMax-H3 inference: #41272
  • [Diffusion] Copy small files into overlay materialized trees instead of linking them: #41298
  • [KDA+Kimi K3] Speed up SANA-Video residual gate add on H200: #41305
  • [Diffusion] Run an all-valid attention mask on the backend's unmasked kernel (Ideogram4 request latency 3.42 to 2.19 s at TP2 on B200): #41309

Local & Desktop AI

  • [SM120] Add optional FlashInfer PCIe-IPC all-reduce for switch-free hosts: #34528
  • fix(moe): support Llama4 NVFP4 router input weights on SM120: #35504
  • [PP][DeepSeek V4] Overlap communication and optimize SM120 prefill: #38792
  • [Fix] Include SM121 in DeepGEMM packed-scale selection: #39482
  • [Qwen4-Exp] Build the offloaded PLE table on the meta device so --ple-offload-embedding never materialises it on the accelerator: #39928
  • Pin triton_kernels num_warps for MXFP4 MoE below Hopper (6x gpt-oss decode on RTX 4090): #41292

AMD / ROCm

  • [AMD] Speed up Wan2.2 DiT FP8 attention per-tensor quantization: #34695
  • [AMD] Fix DSV4 FP4 dequant path for AITER on ROCm: #35123
  • [ROCm][Diffusion] Enable fused qk norm and rope on ROCm: #35573
  • [AMD] Integrate Aiter MegaMoEv2 for DeepSeek-V4: #35619
  • [AMD] Skip full-vocab softmax in EAGLE topk==1 draft on ROCm: #35872
  • [ROCm][Perf] aiter: page-level KV view for gfx950 fp8 page-64 asm prefill: #36505
  • MiniMax-M3: run the sparse prefill main attention through AITER Gluon paged attention: #36546
  • MiniMax-M3: allocate the lightning-indexer K cache in fp8 on gfx95: #36549
  • MoE: small-batch sorting path with fused mxfp8 quantisation: #36559
  • MiniMax-M3: wave64 histogram-select decode top-k, and raise kMaxNumBlocks for CUDA graphs: #36560
  • MiniMax-M3: MXFP8 dense-only block convert + aiter MXFP8 MoE on gfx950: #36574
  • MiniMax-M3: allow shared-experts fusion on ROCm gfx942 and newer: #36576
  • [ROCm] Fix EAGLE spec-decode verify silently sampling greedy on HIP: #37134
  • [ROCm] Widen the HiCache JIT copy rounds and enable the K-only host pool: #37152
  • [AMD] Preserve deterministic inference when Lean Attention is enabled: #37740
  • [AMD][Diffusion] FlyDSL fused norm kernels on wave32 targets (gfx1250): #37751
  • [AMD] Fix DeepSeek-R1-MXFP4 accuracy with AITER FP8: #37762
  • [AMD][DSV4] Enable hicache on deepseek-v4 fp8 unified attn: #37778
  • [ROCm][DSV4] Enable breakable CUDA graph prefill: #37810
  • [AMD] Enable GLM DSA prefill top-k to the v2 kernel (+4.9% token throughput at 70K input on GLM-5.2-MXFP4, 4x MI355X): #37889
  • [AMD][Spec] Enable GDN ReplaySSM target-verify on ROCm: #38184
  • [ROCm] Fuse the MLA q absorb into the RoPE + KV-write kernel on gfx950: #38340
  • [AMD] Avoid the FP8 wo_a path when the weight is BF16: #38453
  • [AMD][GLM-5.3-Flash Day 0] Route mHC through AITER on gfx950: #38545
  • [AMD][GLM-5.3-Flash Day 0] Enable FP8 and Quark MXFP4 MoE on gfx950: #38546
  • [AMD][GLM-5.3-Flash Day 0] Enable zero-RoPE TileLang DSA on gfx950: #38547
  • [AMD][GLM-5.3-Flash Day 0] Honor fused and per-expert names in quark exclude: #39317
  • [AMD][GLM-5.3-Flash Day 0] Enable zero-RoPE MHA prefill on ROCm: #39338
  • [AMD][GLM-5.3-Flash Day 0] Build the fused DSA k-pool top-k JIT kernel on HIP: #39339
  • [AMD][GLM-5.3-Flash Day 0] Support non-2048 top-k widths in the DSA page-table transform: #39340
  • [AMD][GLM-5.3-Flash Day 0] Enable the k-pool DSA indexer on gfx950: #39341
  • [AMD][GLM-5.3-Flash Day 0] Enable speculative decoding (MTP) on ROCm: #39778
  • [AMD][GLM-5.3-Flash Day 0] Load the MXFP4 MTP draft layer: #39779
  • [AMD] Pad QSA MQA decode Q-heads to 16 for ROCm MFMA: #38875
  • [AMD] Add a Triton packed sparse decode path for QSA on ROCm: #38876
  • [AMD] Load fused shared experts for Qwen4-Exp and Qwen3.5 MTP: #38878
  • [AMD][DSV4] feat: enable DSpark with fp8 unified_kv on gfx950: #38901
  • [AMD] Tune Triton sparse MLA on gfx950 and make split-K workspaces graph-safe: #39059
  • [ROCm][Fix] Keep quantization for mixed Quark Qwen3.5 MTP checkpoints: #39064
  • [AMD][Kimi-K3] Fix deferred KDA gate projection and update DCP cookbook: #39066
  • [AMD] GLM-5.2 NextN: cast draft fused MoE to per-channel FP8: #39155
  • [AMD] Use exact CU share for gfx950 segment-plan headroom: #39503
  • [AMD][Fix] Fix vattn_asm HIP error 709 under CUDA graph capture on ROCm 10: #39513
  • [AMD] Fix deferred Kimi-K3 forget gate in fused in-projection: #39525
  • [AMD][Fix] Fix DSV4 MTP crash: #39547
  • [AMD] Prefer HIP Top-K for GLM-5.x on ROCm (P90 interactivity 96.97 to 106.69 tok/s/user on MI355X at concurrency 8): #39631
  • [AMD] Update deepseek-v4 PDI and cache policy setting for agentic workload: #39702
  • [AMD] Clamp MORI intranode grid GPUs: #39763
  • [ROCm] fix: remove extra bf16 -> fp32 cast in jit grouped topk kernel path: #39775
  • [ROCm] feat: enable aiter allreduce fusion for GLM models: #39790
  • [AMD][Fix] Fix dsv4 server launch: #39875
  • [AMD] Reuse KV gather indices across ASM context prefill layers: #39901
  • [AMD] Pack Qwen3.5 GDN input projections on ROCm: #39902
  • [AMD][DSV4] Allow moe_a2a_backend='mori' with DSpark + dp attention: #39910
  • [AMD] Update ROCm AITER pin to acf8fdf9: #39965
  • [AMD] dsv4: pick kv_splits per index stream, not by occupancy alone: #39968
  • [AMD] Use Triton softmax routing for Qwen3.5 on gfx950: #39986
  • [AMD] Tune Qwen3.5 TP4 GDN recurrent launch on gfx950: #39987
  • [AMD][DSV4] fix: drop shadowing local get_exec import that breaks model startup on ROCm: #40186
  • [AMD] Small-M MXFP4 fused-MoE kernel for gfx950 (Qwen): #40204
  • [AMD] Drop the redundant scale zero-fill before AITER per-tensor FP8 quant: #40557
  • [ROCm][DSA] Enable AITER fused FP8 indexer writer: #40710
  • [AMD] Critical fix enabling Qwen3.8 FP8: restore dropped fused shared-expert weights: #40754
  • [AMD] Add diffusion (Wan2.2) extras to gfx1151 Docker image: #40844
  • [AMD][DSV4] fp8 unified_kv decode: wave-aware split count past 40 tokens: #40878
  • [AMD] Drop the unreachable vLLM fallback from ROCm FP8 activation quant: #40879
  • [AMD][DSV4] moe: enable shared-expert fusion on the grouped-topk path (megamoe): #40943
  • [AMD] Add tuned dsv4 shape: #40996
  • [AMD] Add .co for deepseek v4 fp8 decode kernel and add group decode opt: #41120
  • [AMD] Fix int32 offset overflow in Triton DSv4 KV store kernels: #41159
  • [AMD][GLM5] Fuse shared expert into AITER MoE on gfx950: #41161
  • [AMD] Honor an explicit triton moe_runner_backend for mxfp8 on ROCm: #41377

NPU / Ascend

  • feat(hicache): support Ascend Mamba states with FIA and async IO: #32500
  • [NPU] decoding procedure optimization on qwen3.5/3.6: #35958
  • [NPU] Fix xgrammar apply_vocab_mask device dispatch to use torch.ops.npu: #36120
  • [NPU] add chunk gdn kernel and unify ssm state layout for ascend gdn backend: #36187
  • fix(npu): fix hybrid KV transfer with PP prefill in PD disaggregation: #38402
  • [NPU] Refactor weight processing and add NPUSwigluLimit activation: #38420
  • [NPU][Fix] update low latency quantization input and update MXFP8 tests: #38831
  • feat(npu): Support returning indexer top-k results: #39060
  • [NPU][Diffusion] Optimize SenseNova-U1 batched generation: #39382
  • [NPU] Avoid device synchronization in Ascend sampling: #39404
  • [NPU] Support nccl backend for --remote-instance-weight-loader: #39413
  • [NPU] Adapt hicache for K3 hybrid models: #39415
  • dsv4(npu): support prefill context parallelism with interleave and zigzag: #39427
  • [NPU][Diffusion] FA MXFP8 and modelslim w4a4f8 and w8a8f8 support for Wan2.2 and FLUX: #39438
  • [NPU][Fix] Fix undefined require_attn_tp_allgather: #39561
  • [NPU] support kimi k3 on A5 and improve performance: #39589
  • [NPU] Support batch invariant FIA graphs for deterministic inference: #39607
  • [NPU] fix npu docker workspace directory: #39732
  • [NPU][DSV4] dsv4 enable cpp: #39820
  • [NPU] Run arch35 block-FP8 dense linears on the native MXFP8 GEMM: #39823
  • [NPU] Gate DFlash replay metadata refresh behind spec_algorithm check: #39879
  • [NPU] Fuse MXFP4 W4A8 MoE gmm1 + swiglu + requant into one kernel: #39881
  • [NPU] Avoid repeated BF16 wo_a weight transposes in DeepSeek-V4 decode: #39919
  • [NPU][Fix] Avoid M-RoPE recompilation for variable sequence lengths: #40371
  • [Fix][NPU] fix dp-attn hang when pin_mem is True on NPU: #40446
  • [NPU] Fix DSV4 hard-coding kv dtype: #40510
  • [NPU] Update CANN version to 9.1.0: #40524
  • [NPU] support NPU 910C L2 memcache offload: #41527
  • [Docs][NPU] Add MiMo-V2.5-Pro FP4 DFlash best practice on Ascend NPU: #40577

CPU / Intel / XPU

  • [Intel XPU] Enable fused_moe_triton tuning on XPU and add tuned DeepSeek-OCR-2 configs: #28723
  • Speculative Decoding with NGRAM support for XPU: #31362
  • [CPU] Add fp8_per_tensor_scaled_mm_cpu kernel: #32618
  • [XPU] Enable HiSparse hierarchical sparse KV cache on Intel XPU: #32792
  • [CPU][Diffusion] Add fused scale-shift and norm kernels for CPU: #33452
  • [Intel][XPU][KVCanary] Enable KV Canary on Intel XPU: #33520
  • [Intel][XPU] Enable chunked prefill scnearios for XPU with UT: #33804
  • [XPU] Use torch scaled_mm for XPU block FP8 linear: #35605
  • [Diffusion] Fix the XPU capability gates that broke the Wan2.2 A14B DiT path: #36825
  • [XPU] Qwen3.8-flash-next enablement: #37213
  • [CPU] Upgrade torch cpu dependencies to version 2.14.: #37939
  • [XPU] weekly simple model enablement 2026/09/14: #39439
  • [CPU] Add fused_sigmod_mul_cpu operators to the Meta Muse Glimmer model.: #39538
  • [Intel GPU] Xpu/weekly simple model enablement 2026 09 21: #40664
  • [XPU] Support compressed-tensors W4A16 by reusing the torch int4pack path: #40828
  • [XPU] Bump sglang-kernel-xpu wheel to v0.3.0: #41220

Security

  • fix(openai): reject request-supplied chat_template by default: #28135
  • Restrict SafeUnpickler standard-library globals: #39858
  • fix: restrict SafeUnpickler to explicit globals: #40259

Cookbook Updates

  • DeepSeek-V4: a PD-disaggregated arm for B200 FP4 agentic serving with DSpark (#40610); MI355X Pro PD pairs with DSpark and UMBP, plus MegaMoE, FP8 KV attention, and breakable CUDA graph recipes: #41024, #41458
  • GLM-5.2: MI355X MXFP4 recipes move to HIP Top-K and newer images, and the high-throughput cell adds HiCache: #40148, #40570, #41109, #41597
  • GLM-5.3 and GLM-5.3-Flash: reasoning and tool-call parsers on by default, and the DCP option is temporarily removed from GLM-5.3-Flash: #40497, #40036
  • Qwen3.5: GB200 and GB300 configs, and the MI355X HiCache recipe uses the kernel IO backend with the page_first layout: #40770, #39572
  • Qwen3.8-Flash-Next: NVIDIA NVFP4 on B200, B300, and GB300: #41046
  • Kimi-K3: an Ascend 950PR/DT recipe: #40575
  • Intel XPU: the cookbook now marks XPU-supported models: #33649

Dependencies

  • xgrammar 0.2.1 to 0.2.7: #39026
  • cache-dit 1.3.0 to 1.5.1: #40104
  • sentencepiece pinned to 0.2.1 (0.2.2 rejects null pieces in InternVL tokenizers): #41777
  • sgl-eval 0.1.0 to 0.1.2: #40620
  • CUDA 13.4 image: nvidia/cuda:13.4.1-cudnn-devel-ubuntu24.04 base, sgl-kernel 0.4.7, sgl-deep-gemm 0.2.0, sgl-deep-ep 0.1.2: #40987; Triton 3.8 compatibility for DeepSeek-V4.1-Flash on Rubin: #40805
  • CPU: torch 2.14.0, torchvision 0.29.0, triton 3.8.0: #37939
  • XPU: sglang-kernel-xpu 0.2.0 to 0.3.0: #41220
  • ROCm: AITER pin moves to acf8fdf9: #39965
  • NPU: CANN 9.1.0 with Python 3.12 (#40524) and sgl-kernel-npu 2026.9.0.post5: #40157
  • Official releases now also publish a versioned renderer image: #40639

Breaking Changes & Upgrade Notes

  • The unified radix tree uses the Rust core by default. Configurations the Rust core does not support fall back to Python automatically; set SGLANG_UNIFIED_RADIX_TREE_CORE_BACKEND=python to force the Python core everywhere: #39627
  • FlashInfer MoE fused finalize is off by default for numerical accuracy. This costs some W4A4 MoE latency; set SGLANG_FLASHINFER_MOE_FUSED_FINALIZE=1 to restore the fused atomic finalize: #40105
  • PD decode now defers KV release on abort by default on the Mooncake and NIXL backends, holding pages until every prefill rank acks (30 s timeout, SGLANG_DISAGGREGATION_DEFERRED_DECODE_KV_RELEASE_TIMEOUT). Set SGLANG_DISAGGREGATION_DEFERRED_DECODE_KV_RELEASE=False to restore immediate release: #41023
  • Request-supplied chat templates are rejected. A chat_template key inside chat_template_kwargs on /v1/chat/completions or /v1/tokenize now returns an error unless the server starts with --trust-request-chat-template; other chat_template_kwargs are unchanged: #28135
  • Deprecated endpoints, env vars, and aliases are removed. Use /model_info instead of /get_weight_version and /weight_version, /v1/loads instead of /get_load, and POST /hicache/storage-backend/clear instead of /clear_hicache_storage_backend. SGLANG_NSA_* aliases give way to SGLANG_DSA_*, SGLANG_ENABLE_SPEC_V2 is gone, --elastic-ep-rejoin becomes --elastic-ep-join-mode recover, --kv-cache-dtype fp4_e2m1 becomes fp4_mx_block16, and python -m sglang.bench_offline_throughput / sglang.bench_one_batch move to sglang.benchmark.offline_throughput / sglang.benchmark.one_batch: #40795
  • Legacy radix cache implementations are removed. The experimental C++ radix tree (SGLANG_EXPERIMENTAL_CPP_RADIX_TREE) is gone (#40775), and SGLANG_ENABLE_UNIFIED_RADIX_TREE no longer does anything, though it still prints a deprecation warning (#40780). Code importing SWARadixCache, MambaRadixCache, or HiRadixCache must move to UnifiedRadixCache: #40313, #40787
  • Custom logit processors must return full-shape logits. A CustomLogitProcessor that returned a reduced shape such as (1, vocab) and relied on broadcasting now raises at decode time. The built-in processors are unaffected: #39234
  • RL weight updates run inside a session. update_weights_from_tensor and update_weights_from_distributed must be called between begin_weight_update() and end_weight_update(). UpdateWeightsFromTensorReqInput.disable_draft_model is replaced by selector="target", and ChecksumInfo.parallelism_info is now a list with one entry per role: #40777
  • Parallel-state getters are deprecated. They are no longer re-exported from sglang.srt.distributed; read groups through get_parallel() (for example get_parallel().tp_group) or import from parallel_state directly: #40342, #40344, #40707
  • Ascend A2 (910b) images are no longer published. Release and nightly images ship for A3 and 950 (A5) only, on CANN 9.1.0 with Python 3.12: #40524

Don't miss a new sglang release

NewReleases is sending notifications on new releases.