github llm-d/llm-d v0.10.0
Release v0.10.0

latest releases: v0.10, v0
one hour ago

llm-d v0.10.0 Release

Release goal and issues tracked here: #2426, although not all of that was accomplished. Thank you to all our new and old contributors.

High Level Themes

  • Operational hardening — make the production path safe and boring: rollouts, HA, and failure behavior that operators can trust.
  • Production readiness — graduate the flagship paths from "works" to "default."

Core Change Notes

Container build Infrastructure changes

Firstly, as you will see repeated throughout the release notes, llm-d as a community is aiming at moving code upstream, particularly changes regarding maintaining the inferencing images where significant progress has been made this release.

Secondly we got a lot of community feedback regarding our software supply chain and confusing -dev image duplication. The release team has introduced the following image build refactor to address this. Instead of having both -dev and regular images stored on GHCR with different image names, we will instead use quay.io as a staging registry and ghcr.io as our production registry, but container images will have the same names (eg: quay.io/llm-d/llm-d-cuda for dev builds vs ghcr.io/llm-d/llm-d-cuda for production builds).

Any image that is generated from code merging to main or official releases, will exist in the production registry (ghcr.io), and will be signed publicly with cosign for provenance. Any image generated from code on an open pr will be published to quay.io, where images will intentionally NOT BE SIGNED. This is meant to reflect the fact that on PR image builds are a courtesy the llm-d community provides to its developers - helping to facilitate a faster development and validation cycle, but the code's quality, authenticity and integrity cannot be ensured by the llm-d community - run these at your own risk.

NOTE: This is currently ONLY for images produced in llm-d/llm-d (IE the inference server images), and so most notably DOES NOT include the router image and sidecar image. If the community likes the direction we have moved in we will roll this out across the project.

Migrations

  • llm-d/llm-d-kv-cache repo has been migrated to llm-d/llm-d-router repo and deprecation of the old repo is in progress (although is not finished at the time of the llm-d v0.10.0 release).
  • llm-d/llm-d-workload-variant-autoscaler repo has been renamed to llm-d/llm-d-autoscaling to reflect the projects more hollistic approach to autoscaling beyond simply the WVA component. The WVA guides are marked as deprecated and will be removed in a future release.

Deprecations

General note: We are deprecating many of our container image builds now that vLLM supports CI/CD infrastructure to build and publish container images for open PRs. We are making progress towards the same for all accelerators but are committed to keeping our container build infrastructure to support hardware vendors and communities who are in the process of contributing their code upstream.

  • ghcr.io/llm-d/llm-d-cuda container image builds have been DEPRECATED in favor of docker.io/vllm/vllm-openai.
  • ghcr.io/llm-d/llm-d-aws container image builds have been DEPRECATED in favor of public.ecr.aws/deep-learning-containers/vllm.
  • llm-d/llm-d-kv-cache/llmd-fs-connector has been integrated upstream into vLLM under the OffloadingConnector and is thus DEPRECATED. This has the major benefit of existing in tree and out of the box with the container image - no independent installation required.
  • llm-d/llm-d-latency-predictor has been DEPRECATED. The work stream and component has been extremely beneficial, and has contributed greatly to many of the improvements to the various scorers we have seen in the llm-d-router. However we have found the system itself to not give much in the way of performance benefit for the operational complexity that it introduces and so were dropping the workstream for now. We might revisit this in the future and have plans to document and archive our findings in pursing the latencypredictor.

LLM-D v0.9.0 Component Summary

Component Version Previous Version Type
llm-d/llm-d-router-endpoint-picker v0.11.0 v0.10.0 Image + Helm Chart
llm-d/llm-d-router-disagg-sidecar v0.11.0 v0.10.0 Image
llm-d/llm-d-kv-cache - (Moved to llm-d-router) v0.9.0 Library
llm-d/llm-d-inference-sim v0.11.2 v0.10.2 Image
llm-d/llm-d-cuda - (Deprecated) v0.9.0 Image
llm-d/llm-d-aws (EFA) - (Deprecated) v0.9.0 Image
llm-d/llm-d-rocm v0.10.0 v0.9.0 Image
llm-d/llm-d-xpu v0.10.0 v0.9.0 Image
llm-d/llm-d-xpu-sglang v0.10.0 — Image
llm-d/llm-d-cpu v0.10.0 v0.9.1 Image
llm-d/llm-d-kv-cache/llmd-fs-connector - (Deprecated) 0.23 Wheel installed in llm-d
llm-d/llm-d-benchmark v0.8.0 v0.8.0 Image
llm-d/llm-d-workload-variant-autoscaler --> llm-d/llm-d-autoscaling v0.9.0 v0.9.0 Helm Chart + Image
llm-d/llm-d-async v0.10.0 v0.9.0 Helm Chart + Image
llm-d/llm-d-batch-gateway v0.6.0 v0.5.0 Image + Helm Chart
llm-d/llm-d-latency-predictor - (deprecated) v0.9.0 Image
llm-d/mooncake-master-store v0.8.0 v0.8.0 Image
vllm-project/vllm v0.30.0 v0.26.0 Base image
kubernetes-sigs/gateway-api-inference-extension v1.5.0 v1.5.0 Helm Chart + CRDs
kubernetes-sigs/inference-perf v0.7.0 v0.6.1 Tool

Upstream Model Server Images

Engine Image Tag Previous Tag
vLLM docker.io/vllm/vllm-openai v0.30.0 v0.26.0
vLLM Omni docker.io/vllm/vllm-omni v0.28.0 v0.26.0
vLLM TPU docker.io/vllm/vllm-tpu v0.29.0 v0.26.0
vLLM XPU docker.io/vllm/vllm-openai-xpu v0.30.0 v0.26.0
vLLM ROCm docker.io/vllm/vllm-openai-rocm v0.30.0 v0.26.0
vLLM ROCm Omni docker.io/vllm/vllm-omni-rocm v0.28.0 v0.24.1
vLLM CPU docker.io/vllm/vllm-openai-cpu v0.30.0 v0.26.0
llm-d CPU ghrc.io/llm-d/llm-d-cpu v0.10.0 v0.9.0
SGLang docker.io/lmsysorg/sglang v0.5.20 v0.5.16
llm-d SGLang XPU llm-d-xpu-sglang v0.10.0 v0.9.0
TRT-LLM nvcr.io/nvidia/tensorrt-llm/release 1.3.0rc28 1.3.0rc23

Infrastructure Changes

Component Version Previous Version
Gateway API CRDs v1.5.1 v1.5.1
GAIE CRDs v1.5.0 v1.5.0
Istio 1.29.4 1.29.4
AgentGateway v1.5.0 v1.1.0 (as kgateway)
Envoy Gateway / Envoy AI Gateway v1.8.1 / v0.7.0 v1.8.1 / v0.7.0
Leader Worker Set + Dissagregated Set (from: https://github.com/kubernetes-sigs/lws/) 0.11.0 v0.10.0

Release matrix

The llm-d community is committed publishing known and tested software. We have put together a release matrix for this release tracking status of the different guides. Some tests we were not able to run due to infrastructure constraints but it is a strong step towards improving our releases. Please see the v0.10.0 release matrix here: https://github.com/llm-d/llm-d/tree/main/release#release-testing

What's Changed

  • revert artifact upload version bump by @Gregory-Pereira in #2282
  • docs(p2p): apply #2067 review follow-ups by @nilig in #2277
  • fix(ci): stop tiered-prefix-cache GKE GPU nightly lanes from cancelling each other by @sudoalok in #2161
  • Fixing WVA Namespace Overrides: Point the Controller's Watch Target at the Install Namespace - #2253 by @lionelvillard in #2283
  • fix(autoscaling): set WVA --watch-namespace in guide.yaml, not just README by @mamy-CS in #2284
  • [Guides] Update dsv-4 guides by @ilmarkov in #2057
  • reorganize the rollouts docs as a guide by @Gregory-Pereira in #2266
  • Add Grafana dashboard setup documentation for flow control by @alexagriffith in #1629
  • Updating llm-d router metrics across guides by @ahg-g in #2288
  • guides(wide-ep-lws): precise prefix-cache routing variant by @nilig in #2203
  • Coordinator's guide by @roytman in #2119
  • docs: update flow control configuration and observability by @alexagriffith in #2249
  • Fix nightly typos by @Geun-Oh in #2291
  • fix: wrong KEDA_VERSION format in download URL by @weizhoublue in #2290
  • docs(observability): add per-guide troubleshooting for the optimizedbaseline by @gyliu513 in #2129
  • Fix oneCCL init failure for tiered-prefix-cache on Intel XPU by @yuanwu2017 in #2270
  • docs(flow-control): make use-case verification observable and fix guide friction by @LukeAVanDrie in #2213
  • fix(xpu): use llm-d-xpu main image by @xiaojun-zhang in #2289
  • fix(workload-autoscaling): remove obsolete CRD step by @mamy-CS in #2287
  • Organized the batch guides similar to other workload-centric ones by @ahg-g in #2295
  • component bumps by @Gregory-Pereira in #2285
  • ci: restore bespoke flow-control nightly e2e by @LukeAVanDrie in #1918
  • Add variables GATEWAY_API_URL, GAIE_URL and ROUTER_RELEASE_URL automatically calculated from guides/env.sh by @maugustosilva in #2299
  • Update image tags by @diegocastanibm in #2294
  • Harden P2P cache-sharing guide and update GLM results by @nilig in #2296
  • precommit linting by @Gregory-Pereira in #2301
  • Match llm-d.ai/model-id to the served model name by @PrateekKumar1709 in #2310
  • Add multimodal-serving nightlies by @rlakhtakia in #2286
  • Update env for release v0.9.0 by @Gregory-Pereira in #2315
  • Set tags on main branch back to "main" by @maugustosilva in #2316
  • fix CPU builds for NIXL by @Gregory-Pereira in #2314
  • fix: grant FMA Launcher Pods RBAC Permissions to Advertise Serving State by @aavarghese in #2307
  • fix(xpu): restore Wide EP routing sidecar image by @yao531441 in #2319
  • enable swapping between staging and prod registries by @Gregory-Pereira in #2318
  • parrent workflow calls must have permission to write ID tokens by @Gregory-Pereira in #2325
  • docs(router): drop monitoring values from the agentgateway install command by @shimib in #2323
  • [ROCM] bump nightly e2e optimized baseline timeout by @vcave in #2321
  • disable sccache temporarily by @Gregory-Pereira in #2331
  • docs: migrate deprecated plugin and config by @zdtsw in #2320
  • docs(optimized-baseline): fix helm values -f pattern in guide.yaml, the render source by @LukeAVanDrie in #2326
  • remove llm-d aws images by @Gregory-Pereira in #2334
  • Fixes for optimized-baseline guide by @maugustosilva in #2337
  • fix(ci): use env.sh release URL vars in flow-control nightly deploy by @LukeAVanDrie in #2342
  • fix(guides): repair CRD install URLs that break on the latest default by @LukeAVanDrie in #2343
  • fix(templates): replace stale version dropdown, add flow control and scheduling areas by @LukeAVanDrie in #2344
  • docs(well-lit-paths): fix broken observability links in optimized baseline guide by @varad-ahirwadkar in #2364
  • docs(rdma): fix NVSHMEM_DISABLE_GDRCOPY spelling by @245678000000 in #2360
  • docs: update llm-d Router artifact references by @revanthreddy-hai in #2346
  • docs(tiered-prefix-cache): document Intel XPU filesystem tier and use release image by @XinyuYe-Intel in #2312
  • rename default wide-ep to deepseek-r1-0528 by @Gregory-Pereira in #2186
  • feat(guides): migrate flow-control to the guide.yaml method by @LukeAVanDrie in #2345
  • [ROCm] add pd disaggregation AMD ROCm mori-io nightly, split from nixl by @vcave in #2377
  • fix nightly wide-ep-lws workflows after guide directory rename by @zetxqx in #2392
  • Added parameters matrix_type and harness_resources to all CI/CD workflows by @maugustosilva in #2393
  • docs(epp): document saturationDetector nested under flowControl by @rishabhsinha17 in #2383
  • docs(scheduling): remove deprecated pd-profile-handler reference by @roytman in #2376
  • fix(nightly-e2e-verification): reject empty benchmark workspaces by @git-jxj in #2402
  • Release badge matrix based on release branch and nightlies by @diegocastanibm in #2394
  • fix: merge duplicate env blocks in lmcache connector patch by @varad-ahirwadkar in #2373
  • docs: update AKS RDMA deployment guidance by @sulixu in #2397
  • docs: add llm-d router monitoring to remaining well-lit path guides by @0xA8hinav in #2306
  • fix: allowlist retuned in typos config by @MeGaurav4 in #2304
  • Support tiered prefix cache on Intel XPU with vllm native offloading connector with file system by @XinyuYe-Intel in #2239
  • fix(observability): honor selected kubeconfig by @santho090 in #2215
  • Fix sglang image tag v0.5.16.0 -> v0.5.16 by @zetxqx in #2418
  • fix: request 1 DRA claim for SGLang GKE prefill to match TP=1 by @zetxqx in #2420
  • docs(guides): Update Overriding guidance by @jcayab in #2399
  • fix(docker): correct always-true AWS credential check in sccache log by @jcayab in #2413
  • chore(e2e): drop unused OLD_MODEL variable in wide-ep-transform.sh by @jcayab in #2414
  • fix(guides): merge duplicate 'env' key in tiered-prefix-cache LMCache patch by @jcayab in #2415
  • docs(scripts): fix swapped argument order in lint-dockerfile-envvars.py examples by @jcayab in #2417
  • fix(observability): start P/D traffic statistics at zero by @git-jxj in #2403
  • docs(coord-disaggregation): add deployment params and concurrent stress-spike benchmarks by @roytman in #2375
  • guides: make GPU P/D NIXL transport explicit by @nilig in #2371
  • docs(observability): document NIXL transfer metrics by @cyclinder in #2368
  • Add EPP+KEDA+FMA benchmark report (Qwen3-32B / H100) by @dumb0002 in #2317
  • feat(guides): add diffusion-serving text-to-image guide by @zetxqx in #2348
  • docs: drop the InternalRequest envelope from the Pub/Sub quickstart payload by @shimib in #2328
  • Cover prometheus-operated in the Prometheus serving certificate by @PrateekKumar1709 in #2309
  • FMA release upgraded to 0.6.5 and CRDs consolidated by @aavarghese in #2419
  • ci(async-e2e): dump pod diagnostics when the kind e2e fails by @shimib in #2421
  • ci: verify rendered guide READMEs by @yankay in #2384
  • fix: add shared HF cache and disable Xet for vLLM GKE PD overlay by @zetxqx in #2424
  • feat(guides): add diffusion-serving image-to-image guide by @zetxqx in #2350
  • docs: describe flow control for optimized baseline by @nt591 in #2333
  • deps(actions): bump helm/kind-action from 1.14.0 to 1.15.0 by @dependabot[bot] in #2457
  • fix(guides): validate environment variable names by @git-jxj in #2405
  • fix(ci): separate GKE native offloading badge by @git-jxj in #2436
  • fix(interactive-pod): unquote discovered model name by @git-jxj in #2400
  • fix(interactive-pod): scope gateway discovery to namespace by @git-jxj in #2404
  • fix(wva): keep CKS nightly patch out of source tree by @git-jxj in #2452
  • fix(ci): reject HTTP errors in gateway smoke tests by @git-jxj in #2437
  • fix: normalize abbreviated vLLM versions by @git-jxj in #2456
  • fix: handle malformed metrics summaries by @git-jxj in #2450
  • Fix: Remove GKE-specific node selectors from diffusion-serving guides by @weizhoublue in #2430
  • fix(ci): wait for flow-control request bursts by @git-jxj in #2438
  • fix(guides): report unhashable YAML mapping keys by @git-jxj in #2406
  • fix(smoke-test): preserve healthcheck failure accounting by @git-jxj in #2401
  • add high availability preferred backends guide for GKE by @rlakhtakia in #2378
  • fix: correct 'retuned' to 'returned' typo in flow-control/README.md by @Jah-yee in #2305
  • fix(flow-control): validate tuning wizard inputs by @git-jxj in #2455
  • fix: inherit variables between Dockerfile stages by @git-jxj in #2446
  • fix(guides): configure homogeneous TP=2 for SGLang PD disaggregation by @rahulgurnani in #2442
  • add workflow files for ROCm tiered prefix cache by @nitsingh2 in #2463
  • feat(modelexpress-p2p): add gke overlay for A3 Ultra (DRA/RoCE) by @huaxig in #2423
  • feat(snapshot): add GKE fast pod snapshot provider and vLLM wrapper by @RyanRosario in #2105
  • feat(predicted-latency-routing): add Intel XPU support by @yao531441 in #2460
  • tiered-prefix-cache: Add TPU multi-host support and restructure overlays by @dannawang0221 in #2313
  • Update FMA+KEDA benchmark results (Qwen3-32B) and add Qwen3-14B report by @dumb0002 in #2459
  • ci(workload-autoscaling): add AMD ROCm WVA nightly workflows by @PrateekKumar1709 in #2469
  • Expand and update the release matrix by @diegocastanibm in #2471
  • feat(workload-autoscaling): add AMD ROCm support for the WVA nightly by @PrateekKumar1709 in #2467
  • docs: document GKE gateway guide prerequisites by @git-jxj in #2448
  • proposal: add MetaX accelerator support by @FouoF in #2308
  • Temporarily disabled scheduled execution of all OpenShit (IBM) workflows by @maugustosilva in #2477
  • docs: fix ModelExpress base image component path by @git-jxj in #2451
  • docs: add graceful shutdown & request draining ops guide by @azhao155 in #2444
  • docs(observability): add batch-gateway dashboards and alerts to llm-d by @UgaTheDev in #2390
  • docs(README): fix stale version badge and theme count mismatch by @oforiwaasam in #2367
  • Add GKE TPU7x dynamic sub-slicing provider docs and dynamic-slice recipes by @yangligt2 in #2335
  • feat(guides): add diffusion-serving text-to-speech guide by @zetxqx in #2425
  • docs(api): document requestHandler.parsers as a list by @rishabhsinha17 in #2382
  • add guides for ROCm tiered prefix cache (native + lmcache, cpu + fs) by @nitsingh2 in #2464
  • Hardware support: Rebellions NPU (optimized-baseline) by @rebel-minhopark in #2476
  • Update async guides + observability script by @jtechapps in #2468
  • feat: new WLP for KEDA+FMA by @aavarghese in #2191
  • Move FMA+KEDA benchmark reports to the fast-model-actuation-keda WLP by @dumb0002 in #2491
  • docs: add token-aware autoscaling guide (KEDA + EPP token backlog) by @asm582 in #2466
  • [ROCm] Updates for PD disagg mori-io CI by @vcave in #2422
  • ci(wide-ep-lws): add AMD ROCm MoRI wide-EP nightly for DeepSeek-V3 by @shikamd123 in #2494
  • Make DisaggregatedSet the base of the wide-ep guide by @BenjaminBraunDev in #2379
  • feat(guides): add Kueue-based replica rebalancing to workload-autoscaling by @asm582 in #2458
  • docs(guides): remove experimental replica rebalancing by @asm582 in #2498
  • ci(flow-control): increase epp/envoy logging verbosity and capture failed response details by @jtechapps in #2496
  • Hardware support: Iluvatar by @archlitchi in #2381
  • Add Mooncake as a Contributor in ADOPTERS.md by @ykwd in #2504
  • docs: mark Redis Pub/Sub as deprecated in async-processor docs by @804533125 in #2483
  • ci: add past-status workflow for the TensorRT-LLM baseline nightly by @iacker in #2509
  • flow-control: document Intel XPU model server deployment path by @yao531441 in #2507
  • Ci/amd wide ep 2k2k benchmark by @shikamd123 in #2495
  • feat(p2p-kv-cache-sharing): add Intel XPU (TCP-only) variant by @yao531441 in #2506
  • feat: added gke overlay for coordinator by @capri-xiyue in #2503
  • fix(scripts): lint-envvars.py always exited 0, silently disabling the pre-commit/CI gate by @jcayab in #2416
  • docs: add multi-endpoint serving guide by @DeanKelly751 in #2508
  • Guarantee a scale-from-zero path in the fast-model-actuation-keda verification by @dumb0002 in #2492
  • Make DisaggregatedSet the base for the wide-EP DeepSeek-V4 guide by @BenjaminBraunDev in #2501
  • Make DisaggregatedSet the base for the wide-EP GLM-5.2 guide by @BenjaminBraunDev in #2502
  • Update llm-d-async release version to v0.10.0 and use standard EPP saturation metric by @jtechapps in #2517
  • bug: added missing gke overlay and fix tls break change by @capri-xiyue in #2514
  • fix(ci,guides): use GAIE_URL in fast-model-actuation-keda by @jtechapps in #2513
  • docs(wide-ep-lws): link networking prerequisites by @k21993 in #2484
  • docs(wide-ep-lws): order Intel XPU modelserver command before NVIDIA GPU by @yao531441 in #2515
  • docs(autoscaling): make KEDA + EPP the architecture path, deprecate WVA by @lionelvillard in #2497
  • docs: migrate guides from removed router plugins by @zdtsw in #2499
  • docs(guides): add GKE Pod Snapshots single-GPU user guide by @RyanRosario in #2489
  • Fix: Wrong URL to Download Latest GAIE Manifests by @weizhoublue in #2527
  • Bump up sglang version to the 0.5.19 version by @rahulgurnani in #2526
  • Rename the FMA guide directory to `fast-model-actuation-base by @dumb0002 in #2521
  • docs(guides): rename gke-pod-snapshots guide to pod-snapshot by @RyanRosario in #2525
  • docs(autoscaling): fix EPP image pin value by @nicole-lihui in #2539
  • docs(observability): fix EPP tracing values key by @nicole-lihui in #2538
  • fix(guides): make calibration script executable by @tangming1996 in #2543
  • fix(workload-autoscaling): avoid mutating AMD model patch by @tangming1996 in #2541
  • Added recipe for inference cost by @simanadler in #2510
  • fix(interactive-pod): respect namespace for gateway lookup by @tangming1996 in #2545
  • feat(sglang): add tiered prefix cache GKE nightly lane by @tyuchn in #2480
  • Rename the wide-ep-lws guide to wide-ep by @BenjaminBraunDev in #2512
  • fix(p2p): enable experimental router plugins by @yankay in #2548
  • Re-enabling automation for workflow by @aavarghese in #2549
  • fix(smoke-test): rebuild payload for chat API fallback by @tangming1996 in #2542
  • fix(ci): unpin SGLang tiered cache nightly cluster by @tyuchn in #2547
  • [Docs] Replace deprecated kv_both for NIXLConnector and with explicit P/D Roles by @NickLucche in #1715
  • fix(e-disaggregation): drop placeholder image from e-pd encode patch by @capri-xiyue in #2559
  • Update multimodal guides to use leaf name for ci/cd testing.- #2556 by @capri-xiyue in #2561
  • Update multimodal guides to use leaf name for ci/cd testing. by @rlakhtakia in #2556
  • fix(sglang): align P/D KV-page size with EPP chunking and sibling guides by @tyuchn in #2500
  • fix(pd-disaggregation): add shared HF cache to vLLM CoreWeave overlay by @pujitha24 in #2528
  • docs(autoscaling): deprecate WVA and pin its image to release-0.9 by @lionelvillard in #2490
  • docs: add a model loading and startup acceleration guide by @yankay in #2540
  • pd-disaggregation: fold the TPU guide into the main guide, cache TPU weights on the node, pin TPU7x to vllm-tpu v0.26.0 by @yangligt2 in #2533
  • fix(pd-disaggregation): drop the ROCM_ATTN backend pin from the AMD overlay by @BenjaminBraunDev in #2558
  • fix(autoscaling): use plain HTTP for keda-epp generic-k8s prometheus scaler by @mamy-CS in #2571
  • Delete WVA nightly E2E CI by @asm582 in #2576
  • refactor(autoscaling): consolidate keda-epp queue and saturation guides by @mamy-CS in #2577
  • Enable cron schedule for multimodal tests by @rlakhtakia in #2566
  • fix: resolve typos linter false positives for AIMD and opt-in by @varad-ahirwadkar in #2551
  • Add COMPONENTS.md: core and ecosystem component roles by @chcost in #2553
  • feat(observability): add GKE TPU dashboard by @bzsuni in #2487
  • docs(tiered-prefix-cache): add observability and troubleshooting section by @k21993 in #2485
  • Slack notification workflow by @diegocastanibm in #2366
  • Guides definition by @diegocastanibm in #2349
  • CI - strip out llm-d cuda images + deepep v2 by @Gregory-Pereira in #2339
  • chore(metrics): old metrics have been dropped in 0.11.0 router by @zdtsw in #2586
  • fix: missed rename to fma-base causing nightly to fail by @aavarghese in #2585
  • [Release] Add a "prepare_for_release.sh" helper by @maugustosilva in #2589
  • docs(release): release testing matrix for v0.10.0 by @github-actions[bot] in #2590
  • pass dry-run to dispatched lanes in release-e2e by @diegocastanibm in #2591
  • Move GKE PD runs to a new cluster by @achandrasekar in #2593
  • Set UCX_IB_ROCE_REACHABILITY_MODE=all in the wide-ep GKE overlay by @BenjaminBraunDev in #2562
  • Bump guides to LWS v0.11.0 + shorten naming by @BenjaminBraunDev in #2578
  • never-run badges by @diegocastanibm in #2594
  • chore: route nightly alerts to ci channel by @zdtsw in #2598
  • docs: fix broken UCCL repository link by @lizzzcai in #2604
  • Preparation for release v0.10: updated all model server engines by @maugustosilva in #2595
  • Bump versions release 0.10 by @Gregory-Pereira in #2596
  • docs: fix broken relative links by @AlSh007 in #2599
  • docs(workload-autoscaling): explain EPP flow-control on/off for keda-epp thresholds by @mamy-CS in #2605
  • Fix GKE Wide-EP guide and nightly E2E workflow by @achandrasekar in #2602
  • feat(workload-autoscaling): default keda-epp startup-time mitigation via HPA stabilization windows by @mamy-CS in #2607
  • Update accelerator specific images (ROCM and XPU) to v0.10.0 by @maugustosilva in #2608
  • fix: Enable vLLM render endpoints for cache routing by @xiaojun-zhang in #2609
  • Fixed an issue with vllm version 0.30 and predicted-latency-routing by @maugustosilva in #2611

New Contributors

Full Changelog: v0.9...v0.10.0

Don't miss a new llm-d release

NewReleases is sending notifications on new releases.