Changes since v0.20.0:
Actions Required Before Upgrading
(No, really, you MUST read this before you upgrade)
- Minor releases: Review the
.0release notes for each new minor version you cross; see:v0.19.0,v0.20.0.
Changes by Kind
Feature
- ElasticJobsViaWorkloadSlices: Fixed a bug where elastic Pods ungated for a ProvisioningRequest could carry the consume-provisioning-request annotation without provisioning-class-name when they were created before the job template was stamped. Kueue now applies both annotations, preventing such Pods from being pinned by the cluster autoscaler to nodes that are never created. (#16543, @neilb-dotcom)
Bug or Regression
- CLI: Fixed a bug where
kueuectl create clusterqueuefailed with "must match the name in coveredResources" when flavors listed the same resources in a different order, for example--nominal-quota "alpha:cpu=1;memory=1,beta:memory=2;cpu=2". (#16544, @Eason09053360) - CLI: Fixed a bug where
kueuectl delete workloaddeleted the owning batch/v1 Job but left its Pods running and the Workload in place. Owners are now deleted with background propagation, matchingkubectl delete. (#16511, @Eason09053360) - ElasticJobsViaWorkloadSlices: Fixed a bug where a replacement Workload could remain pending when one PodSet shrank below its predecessor's grant while another grew. (#16892, @weizhoublue)
- ElasticJobsViaWorkloadSlices: Fixed a bug where resizing an elastic job could let the scheduler admit workloads beyond a cohort's quota in the same scheduling cycle. With ElasticJobsViaWorkloadSlicesWithTAS enabled, the replaced slice's topology usage no longer defers other workloads to the next scheduling cycle. (#16601, @apullo777)
- FairSharing: Fixed
kueue_cohort_weighted_shareto reportNaNinstead of+Infwhen a zero-weight Cohort is borrowing quota, matching its help text andkueue_cluster_queue_weighted_share. (#16818, @abhayjoshi201) - FairSharing: Fixed a bug where a ClusterQueue or Cohort whose weighted share exceeded the int64 range could report a negative
status.fairSharing.weightedShare. Such a share is now reported as the int64 maximum. (#16592, @thc1006) - FairSharing: Fixed a bug where a ClusterQueue or Cohort with weight 0 that borrowed only resources its Cohort had nothing lendable for got a weighted share of 0 instead of the maximum, and was ranked as if it were within its quota. (#16567, @thc1006)
- FairSharing: Fixed a bug where a borrowing ClusterQueue or Cohort with a very large weight, such as 1e308, could report a
status.fairSharing.weightedShareof 0 and be ranked as if it were within its quota. Kueue now keeps such a share above zero: the status reports 1, and thekueue_cluster_queue_weighted_shareandkueue_cohort_weighted_sharemetrics report the smallest positive float64 value. (#16584, @thc1006) - FairSharing: Fixed a bug where restoring same-
ClusterQueueworkloads could make cross-ClusterQueuepreemption targets violate fair-sharing rules and cause preemption loops. Kueue now validates targets against the final simulated post-admission state. (#16912, @JanKaczmarski) - JobFramework: Fixed Workloads inheriting Kueue's internal labels and annotations from the Job or Pod when they were listed in labelKeysToCopy or as custom metric sources. (#16773, @thc1006)
- KueueViz: Fix the LocalQueue detail page missing the Status cell in the workloads table and returning null instead of an empty array for queues with no workloads. (#16692, @Dasmat13)
- KueueViz: Fix the workload detail page not showing the ClusterQueue name and preemption state. (#16608, @onkar717)
- KueueViz: Fixed a bug where the workloads dashboard listed unrelated Pods for Workloads without a
kueue.x-k8s.io/job-uidlabel, such as pod groups. Such Workloads now show an empty Pods list. (#16702, @henry3260) - MultiKueue: Enabling MultiKueueRayInTreeAutoscaling now requires SkipChildJobSuspension to be enabled. Kueue reports a startup validation error if this dependency is disabled or omitted. (#16814, @kevin85421)
- MultiKueue: Fixed a bug where a replaced workload slice could keep holding quota on a worker cluster when the scheduler failed to finish it after admitting the replacement slice. (#16772, @kevin85421)
- Observability: Fix stale pending-workload metrics when LocalQueues are created, resumed, stopped, or deleted, including Workloads observed before their LocalQueue. (#16890, @tomsen02)
- Observability: Fixed a bug where the
kueue_replaced_workload_slices_totalmetric was not exposed on the metrics endpoint. (#16763, @Narwhal-fish) - Observability: Fixed stale
kueue_pending_workloadsseries that remained after a ClusterQueue was stopped or resumed whenCustomMetricLabelswas enabled with a Workload-sourced custom label, which caused pending workloads to be double-counted. (#16871, @kevin85421) - Observability: Refresh pending-workload metrics when ClusterQueues stop or resume, without relying on subsequent Workload updates. (#16891, @tomsen02)
- Ray integrations: Fixed a bug where workloads with worker groups using replicas and numOfHosts were not admitted because SubGroupCount and topology index labels were missing from the topology request. The KubeRayEvictOnInconsistentTopologyRequest feature gate is disabled by default, allowing existing workloads to continue running with old topology requests. (#15677, @spencer-p)
- Ray integrations: Use zero instead of one as the worker replica fallback when replicas is absent, matching the KubeRay CRD default. (#16908, @kevin85421)
- RayService: Added the Alpha
KubeRayServiceUsingTopLevelSuspendfeature gate, disabled by default, to opt in to top-levelspec.suspendhandling. Upgrade KubeRay to v1.7.0 or newer before enabling the gate. (#16882, @kevin85421) - Scheduling: Fixed a bug where Workloads blocked by a closed preemption gate could retain a stale
QuotaReservedcondition andstatus.resourceRequests. Such Workloads now report thePreemptionGatedreason on theQuotaReservedcondition instead ofAdmissionGated. (#16699, @henry3260) - Scheduling: Fixed a bug where a CPU amount past the Quantity limit appeared in quota messages without the milli suffix, so it read as cores. (#16716, @HarshithaMS005)
- Scheduling: Fixed a bug where a Workload whose total request exceeded the int64 range could be admitted against a quota it did not fit. (#16549, @HarshithaMS005)
- Scheduling: Fixed a bug where partial admission of a PodSet whose total request overflowed int64 could admit more Pods than the quota allows, because the reduced request was derived from the saturated total instead of the per-Pod request. (#16519, @thc1006)
- StatefulSet: Fixed a bug where changing a StatefulSet’s queue while its Workload held quota could cause repeated Workload update failures when the StatefulSet had no ready replicas. Kueue now rejects the queue change until the Workload releases its quota. (#16705, @weizhoublue)
- StatefulSet: Fixed a bug where labels listed in
integrations.labelKeysToCopy, and the labels and annotations used asCustomMetricLabelssources, were not copied from a StatefulSet into its Workload. (#16751, @nsega) - TAS: Fixed a bug where a Workload using delayed topology assignment or failed node replacement could lose its quota reservation, and for a failed node replacement also its admission. (#16491, @apullo777)
- TAS: Fixed a bug where a Workload using delayed topology assignment or failed node replacement could miss a scheduling retry until it was updated again. (#16503, @apullo777)
- TAS: Fixed a bug where a second scheduling pass (failed-node replacement or delayed topology assignment) computed from an outdated Workload could overwrite a newer admission or drop status updated in the meantime, such as an Evicted condition. The scheduler now rejects the stale write, keeps the Workload's reservation, emits an OutdatedScheduleCycle event and retries the second pass. (#16858, @anguszzzz)
- TAS: Fixed a bug where an elastic Workload with a TAS PodSet group could reassign its running leader to a different domain when the workers scaled up from zero, with the
ElasticJobsViaWorkloadSlicesWithTASfeature gate enabled. (#16820, @Narwhal-fish) - TAS: Fixed a bug where stuck pods could break rank-based pod ungating after node replacement or workload re-admission. (#16796, @pajakd)
- TAS: Fixed a scheduler panic when an elastic Workload scaled the workers of a PodSet group from 0 to the same count as the leader, with the
ElasticJobsViaWorkloadSlicesWithTASfeature gate enabled. (#16810, @Narwhal-fish) - TAS: Fixed workloads with non-adjacent PodSet group members getting stuck during flavor-scan resume or delayed/failed-node topology reassignment. (#16757, @apullo777)
- TAS: With
KueueDRADeviceFeasibilityenabled, fixed a bug where PodSets of one Workload could be placed on a node whose device plugin publishes fewer devices of a DRA-backed extended resource than they need together, leaving Pods Pending. The same miscount could keep an elastic scale up pending. (#16809, @thc1006)