github ray-project/ray ray-2.59.0
Ray-2.59.0

10 hours ago

Highlights

  • Ray Data LLM & Ray Serve LLM are GA/Stable: the LLM APIs graduate to general availability this release (#65194), alongside an upgrade to vLLM 0.27.0 (#65351).
  • Ray Data external shuffle (shuffle_v2): the new shuffle backend gains an external, disk-backed mode — a new external shuffle runtime library and task/operator set (#64828, #65144), planner integration (#65499), and Join & Aggregation support on external shuffle (#65897). The strategy was renamed from hash_shuffle_v2 to shuffle_v2 (#65411) and the compression setting from hash_shuffle_compression to shuffle_compression (#65590).
  • Token authentication: Ray now turns on token authentication by default for local clusters. ray.init() without an address enables authentication and generates a token at ~/.ray/auth_token when none exists. ray start --head enables authentication when a token is already available, and otherwise warns that it started an unauthenticated cluster. Remote and multi-node clusters are unchanged in this release. Set RAY_AUTH_MODE=disabled to opt out (#64755).
  • Security hardening: the dashboard redacts runtime_env values (which routinely carry credentials) in browser-facing endpoints (#65226), and read_hudi gained an unpickling guard preventing remote code execution (#65780).
  • Ray Core at scale: removed the GCS resource-view reset that caused placement-group retry storms on large clusters (#65271), hardened the worker-lease protocol (at most one CancelWorkerLease per pending lease, lease-id tombstones — #65613, #65420), and batched local object frees (#65000, #64037).
  • New Data sources: native ORC file reading (#64600), multi-path read_lance (#64560), and fps/resize parameters for read_videos (#64120).

Ray Data

🎉 New Features

  • Add support for reading ORC files (#64600)
  • Introduce external disk shuffle: runtime library, tasks and operators, planner integration, and Join & Aggregations on external shuffle (#64828, #65144, #65499, #65897)
  • Support reading multiple paths in read_lance (#64560)
  • Add fps and resize parameters to read_videos() (#64120)
  • Add max_concurrent_calls_per_actor to ActorPoolStrategy (#65534)
  • Introduce reporting for task USS distribution metrics (#65666)

💫 Enhancements

  • Iceberg: deduplicate read-task state, better decoded-size estimation, and bounded read-task memory (#64811, #65042, #64813)
  • Count materialized data as downstream capacity, and return external consumer bytes for downstream-capacity backpressure (#66226, #65474)
  • Disable throttling for Union and Mix operators (#65242)
  • Pin the fused map function for shuffle tasks (#65480)
  • Chunk DataFrame emission in WebDatasetDatasource._read_stream and detect non-contiguous keys in read_webdataset (#65394, #63407)
  • Expose read-task resource args in read_webdataset (#65623)
  • Auto-detect AzureFileSystem for Blob HTTPS URLs in download() (#63648)
  • Groundwork for footer-based Parquet reads: footer types and online bin packer, FooterReader actor pool, and ListFiles pushdown state derived from the ReadFiles scanner (#65210, #65273, #65214)
  • Groundwork for Data fault tolerance (1/n): add a linear application-layer lineage tracker (#65519)

🔨 Fixes

  • Add an unpickling guard to prevent RCE when reading Hudi (#65780)
  • Fix Iceberg Dataset.count() returning wrong row counts (#65175)
  • Fix LeRobot delta windows for multidimensional features (#65165)
  • OpTask._cancel never passed force=True (#65389)
  • Fix an incorrect timeout value in the autoscaler shutdown warning (#65317)
  • Correct the physical-to-logical operator mapping type (#61790)
  • Don't register out-of-scope callbacks on blocks passed through from input to output (#66369)
  • Fix crash in preprocessors due to stale metadata in Arrow-backed Pandas dataframes (#66578)

📖 Documentation

  • Style and accuracy passes on the shuffle, join/save, and OOM-prevention pages (#65374, #65372, #65525, #65509)
  • Clarify num_workers in the PyTorch docs and iter_torch_batches collate_fn usage (#65311, #64535)
  • Deprecations: ray_remote_args for read APIs (#65501) and for Dataset transformations (#65228); min_rows_per_file with partitioned Parquet writes (#63368)
  • Rename the shuffle strategy from hash_shuffle_v2 to shuffle_v2 (#65411) and hash_shuffle_compression to shuffle_compression (#65590)

Ray Serve

🎉 New Features:

  • Configurable backpressure responses. A new per-deployment BackpressureConfig lets you return 429 instead of 503 when a request is rejected for exceeding max_queued_requests, and optionally attach a Retry-After header. This separates deliberate load shedding from "the deployment is broken" for load balancers, retry policies, and availability SLOs, and brings the HTTP path to parity with gRPC's RESOURCE_EXHAUSTED. Defaults are unchanged (#65193).
  • Declarative tracing configuration. A new TracingConfig model (enabled, exporter_import_path, sampling_ratio) can be passed to serve.start(tracing_config=...) or set in a Serve config file, replacing environment-variable-only setup and wiring tracing through the proxy and replica paths. Setup errors now fail fast instead of being silently swallowed (#63273).
  • Scale-to-zero for gang-scheduled deployments. Gang-scheduled deployments can now set min_replicas=0; the autoscaler preserves the final 1 → 0 transition instead of rounding it back up to a full gang (#65575).

💫 Enhancements:

  • Faster unary gRPC on direct ingress. Unary calls now take a dedicated fast path instead of being funneled through an async generator, recovering most of the throughput lost when the handlers were unified (~4684 vs. 4601 avg RPS in the release benchmark) (#65398).
  • Application builds respect node label selectors. The build task now inherits a label selector derived from deployment actor options when all deployments agree on one, so the application import runs on a node pool with the right image instead of landing on an incompatible head node (#65180).
  • Layer-4 mark-down for HAProxy servers, behind an environment-variable gate, so draining backends are observed at the connection level (#65267).
  • Token authentication is on by default for local clusters, and Serve's client and API calls propagate the token end to end (#64755, #66248).
  • Lighter imports. jinja2 is now imported lazily and declared in the serve extra, so it is no longer pulled in on every Serve import (#65542, #65988).
  • Deprecated APIs removed. deploy_mode / ServeDeployMode, use_new_handle_api on DeploymentHandle.options, and the RAY_AGENT_ADDRESS deprecation path are gone, and a # in a deployment name is now a hard ValueError since it is the replica ID delimiter (#65215).
  • Explicit public surface for request routers. ray.serve.request_router now declares __all__, making the custom-router extension API visible to the API consistency checks (#65485).

🔨 Fixes:

  • Controller no longer crashes during RayService incremental upgrades. Shutting down a proxy pinned to a node that no longer exists raised an uncaught ActorUnschedulableError and killed the controller loop; the handle is now cleaned up gracefully. This showed up with GCS fault tolerance backed by Redis, where the new cluster restores proxy actors pinned to pre-upgrade node IDs (#65076).
  • Global tracing config no longer leaks between tests. A config test left tracing enabled at sampling ratio 1.0 for every subsequent test in the session (#65432).
  • Controller benchmark stability at scale. The sweep now extends to 8K replicas with a flat peak-CPU budget and a larger head node, and reserves per-replica memory to stop OOM kills at the 8K checkpoint (#65371, #65847, #65987).
  • Test suite deflaking across the CLI, GCS failure, cluster, proxy, gRPC, cancellation, and crashed-replica tests, mostly by replacing fixed request counts and inherited 10s timeouts with explicit deadlines (#65464, #65488, #65537, #65560, #65617, #65619, #65624).

📖 Documentation:

  • Documented the # restriction on deployment names and why it exists, so the new ValueError has something to point at (#65489).
  • Added a KV-aware routing guide covering installation, configuration, scoring, and tuning, and extended the KV cache offloading guide with the native vLLM backend and new Grafana panels (#65569).
  • Added a guide for serving LLMs on TPUs (#65026).
  • Docs infrastructure: curated descriptions for every page in llms.txt, high-traffic landing pages and card grids converted from RST to MyST, explicit image versions in place of floating latest tags, and the "Ray Pod" term retired from the Kubernetes docs (#65115, #65467, #65472, #65339, #65423).

🏗 Architecture refactoring:

  • Static type checking enforced across Serve's core. mypy and pyrefly now run on the controller-critical modules (controller, deployment_state, replica, application_state), the proxy and data-plane modules (proxy, proxy_state, haproxy), and the scheduling and routing modules (deployment_scheduler, router, request_router). Beyond annotations, this tightened several honest-return-type and ReplicaID-vs-str key confusions in state-machine code where type confusion becomes persistent state corruption (#64753, #65332, #65410).

Ray Data LLM / Ray Serve LLM

🎉 New Features

  • Ray Data LLM and Ray Serve LLM are GA (#65194)
  • Upgrade to vLLM 0.27.0 (#65351).
  • Support multipart/form-data (file uploads) in HttpRequestUDF (#63903)

💫 Enhancements

  • Defer heavy engine imports from ray.data.llm (#64990)
  • Avoid vLLM dtype auto-detection in multimodal prep (#64015)
  • Surface engine errors on the direct-streaming ASGI app (#65440)

🔨 Fixes

  • Fix s3:// model_source being dropped for streaming load formats (#65753)

📖 Documentation

  • Add a KV-aware routing guide and native KV-cache offloading docs (#65569); add docs for serving LLMs on TPUs (#65026)

Ray Core

🎉 New Features

  • Enable token authentication by default for local clusters. ray.init() generates and reuses a token automatically; ray start --head enables authentication when a token is available from RAY_AUTH_TOKEN, RAY_AUTH_TOKEN_PATH, or ~/.ray/auth_token. RAY_AUTH_MODE=disabled opts out (#64755)
  • Add opt-in swap accounting to the memory monitor and scheduler (#63793)
  • RDT/NIXL: enable driver-side ray.put with NIXL tensor transport, and let the pool serve tensors on a different device (#65072, #65418)
  • Mobilint accelerator support (#61898)
  • TPU: add subslice_index to subslice_placement_group (#65403)
  • Add opt-in UNLINK-based Redis namespace cleanup (#65522)
  • Add RAY_DISABLE_WORKER_LOG_PREFIX to control log prefix behavior (#65230)
  • GCS Active-Passive phase 2.1: leader-election interface (protocol, status, client cache) (#65132)

💫 Enhancements

  • Remove the GCS resource-view reset that caused placement-group retry storms at scale (#65271)
  • Guarantee each pending lease request receives at most one CancelWorkerLease RPC, and tombstone lease ids on cancellation (#65613, #65420)
  • Batch local object frees and gut internal free paths (FreeObjects series) (#65000, #64037)
  • Free unconsumed objects reported for deleted generators (#65276)
  • Avoid per-call map allocation in std::hash<ResourceSet> (#64958)
  • Assert pub-sub channel subscription counts to catch unexpected growth (#65148)
  • Sandbox: preserve archived mtimes when extracting image layers (#65737)

🔨 Fixes

  • Fail loudly on an unhandled GetObjectStatus reply instead of hanging (#65270)
  • Fix GetLocationFromOwner treating timeout_ms as microseconds (#65260)
  • Fix None yield from a restarted streaming generator with application errors (#65121)
  • Clean up async ObjectRefGenerator timeout waiters (#65649)
  • Enforce the task_id/put_index contract in GetGeneratorReturnId (#65301)
  • Use forkserver for the POSIX multiprocessing context (dashboard) (#65159)

Ray Clusters & Autoscaler

💫 Enhancements

  • Propagate worker-group priority from the RayCluster CR to the autoscaler (KubeRay) (#65244)
  • Route autoscaler INFO logs to stdout instead of stderr on KubeRay (#65455)
  • Support configurable Kubernetes API authentication in the autoscaler (KubeRay) (#65827)

🔨 Fixes

  • Retry AWS key-pair creation after a duplicate error (#64738)
  • Deduplicate cloud instances during termination (#65419)
  • Fix instance starvation from request-ID sharing (launch_errors dict collision and request_id reuse) (#65299)
  • Skip DEAD GCS nodes in LoadMetrics so they cannot overwrite a live node's idle state by IP (autoscaler v1) (#65592)

Dashboard & Observability

  • Redact runtime_env values in the dashboard (#65226)
  • Configurable defaults and UI dialogs for py-spy/memray profiling parameters (#64806)
  • Fix component_cpu_percentage always reporting zero for dashboard subprocess modules (#65055)
  • Support events.k8s.io/v1 reportingComponent fallback in the KubernetesEventProvider (#65192)

Ray Train

  • Update elastic-training accelerator examples in the docs (#64284)
  • New release-test benchmarks: preemption benchmark and a Qwen3-0.6B DeepSpeed benchmark harness (#65047, #64315); pin train and validation benchmark workers to their subclusters (#66041)

RLlib

  • Fix APPO hyperparameter tuning (#65528)
  • Type from_checkpoint return as Self (#65645)
  • Rebuild the RLlib release-test learning suite (#64781)

Java

  • Update httpcore5 to 5.4.3 (#66081)

Documentation

  • Convert the highest-traffic landing pages, library front doors, and remaining card grids from RST to MyST (#65467, #65470, #65472)
  • Add an in-repo llms.txt generator (replacing sphinx-llms-txt), curated page descriptions, and a note that doc pages also serve Markdown for agents (#64458, #65115, #65461)
  • KubeRay docs wave: NetworkPolicy user guide, Kubernetes/Ray scheduling orientation guide, SidecarSubmitterRestart guide, built-in Ingress ingressOptions, KubeRay v1.7 reference updates, and a vendored CRD API reference (#65074, #65263, #64303, #65483, #65498, #65513, #65428, #65618)
  • User guide for JAX TPU profiling on GKE, with a style pass (#64735, #65287)
  • Render API type hints in both the signature and the Parameters list; generate heading anchors for h4 headings (#64778, #65240)
  • Deprecate the ray-ml images in the docs and stop recommending them; pin explicit Ray image versions in place of floating tags (#65388, #65339)
  • Prevent dismissed announcement banners from flashing on page load (#66343)

Build & Dependencies

  • Downgrade protobuf to 5.27.5 (#66119)
  • Pin the uv installer version in release and CI image builds (#66068)
  • Build the manylinux wheels without pip build isolation (#65577)
  • CI now resolves Python packages and bazel downloads through a CI-hosted mirror (#65687, #65571, #65599, #65605)
  • Ray 2.59.0 is the last release that will publish CUDA 11.7 (-cu117) images.

Upcoming

  • Upcoming in Ray 2.61: Token authentication becomes the default for all clusters, including remote and multi-node clusters. Every node in a cluster and every client that connects to it will need the same token. Operators of remote clusters should set up token distribution before upgrading. See Ray token authentication for how to generate and distribute a token.

Thanks

Many thanks to all those who contributed to this release!

@mikemikimike, @ShockYoungCHN, @elliot-barn, @aaronlinear, @spencer-p, @nadongjun, @marwan116, @Kunchd, @avigyabb, @suppagoddo, @harshit-anyscale, @n3sfan, @Smallfu666, @bveeramani, @pjdurden, @Jenson97, @gangli113, @ronny-anyscale, @wufeiye, @preneond, @ankushbbbr, @yuhuan130, @ixsanpe, @dossett, @rayhhome, @ArturNiederfahrenhorst, @robertnishihara, @CrossManger, @hogeheer499-commits, @iamjustinhsu, @justinvyu, @wuallen57730, @dstrodtman, @dragongu, @ryanaoleary, @goutamvenkat-anyscale, @YashwanthRanjanSingaravel, @400Ping, @rueian, @imtherealnaska, @Sparks0219, @NightWing1998, @MortalHappiness, @lorriexingfang, @pseudo-rnd-thoughts, @aliceco01, @sai-miduthuri, @leauny, @LuciferYang, @Aydin-ab, @neidythedev, @hsusul, @richabanker, @eicherseiji, @thomasdesr, @justinyeh1995, @Hyunoh-Yeo, @owenowenisme, @neuyilan, @vinay7373, @vivekmahajan, @andrewsykim, @zzchun, @johntaylor-cell, @liulehui, @Yicheng-Lu-llll, @khluu, @machichima, @xyuzh, @chipspeak, @dataminsu, @AarryaSaraf, @YoyinZyc, @chiayi, @viiccwen, @nightcityblade, @lonexreb, @jeffreywang88, @nikitagrover19, @sampan-s-nayak, @onlinerj, @abligail, @karticam, @ayushk7102, @win5923, @xubo245, @BatshevaBlack, @iacker

Don't miss a new ray release

NewReleases is sending notifications on new releases.