github kubernetes-sigs/kueue v0.19.8

latest release: v0.20.1
5 hours ago

Changes since v0.19.7:

Actions Required Before Upgrading

(No, really, you MUST read this before you upgrade)

Changes by Kind

Bug or Regression

  • CLI: Fixed a bug where kueuectl create clusterqueue failed with "must match the name in coveredResources" when flavors listed the same resources in a different order, for example --nominal-quota "alpha:cpu=1;memory=1,beta:memory=2;cpu=2". (#16545, @Eason09053360)
  • CLI: Fixed a bug where kueuectl delete workload deleted the owning batch/v1 Job but left its Pods running and the Workload in place. Owners are now deleted with background propagation, matching kubectl delete. (#16512, @Eason09053360)
  • ElasticJobsViaWorkloadSlices: Fixed a bug where resizing an elastic job could let the scheduler admit workloads beyond a cohort's quota in the same scheduling cycle. With ElasticJobsViaWorkloadSlicesWithTAS enabled, the replaced slice's topology usage no longer defers other workloads to the next scheduling cycle. (#16125, @apullo777)
  • FairSharing: Fix kueue_cohort_weighted_share metric to report NaN instead of +Inf when a Cohort with fairSharing.weight: 0 is actively borrowing resources, matching kueue_cluster_queue_weighted_share and the metric documentation. (#16864, @abhayjoshi201)
  • FairSharing: Fixed a bug where a ClusterQueue or Cohort whose weighted share exceeded the int64 range could report a negative status.fairSharing.weightedShare. Such a share is now reported as the int64 maximum. (#16593, @thc1006)
  • FairSharing: Fixed a bug where a ClusterQueue or Cohort with weight 0 that borrowed only resources its Cohort had nothing lendable for got a weighted share of 0 instead of the maximum, and was ranked as if it were within its quota. (#16568, @thc1006)
  • FairSharing: Fixed a bug where a borrowing ClusterQueue or Cohort with a very large weight, such as 1e308, could report a status.fairSharing.weightedShare of 0 and be ranked as if it were within its quota. Kueue now keeps such a share above zero: the status reports 1, and the kueue_cluster_queue_weighted_share and kueue_cohort_weighted_share metrics report the smallest positive float64 value. (#16587, @thc1006)
  • FairSharing: Fixed a bug where restoring same-ClusterQueue workloads could make cross-ClusterQueue preemption targets violate fair-sharing rules and cause preemption loops. Kueue now validates targets against the final simulated post-admission state. (#16911, @JanKaczmarski)
  • JobFramework: Added the Alpha SkipChildJobSuspension feature gate (disabled by default). When enabled, Kueue no longer suspends managed child jobs based on the ancestor Workload's state and leaves their lifecycle to the ancestor controller. (#16460, @kevin85421)
  • JobFramework: Fixed Workloads inheriting Kueue's internal labels and annotations from the Job or Pod when they were listed in labelKeysToCopy or as custom metric sources. (#16774, @thc1006)
  • KueueViz: Fix the LocalQueue detail page missing the Status cell in the workloads table and returning null instead of an empty array for queues with no workloads. (#16693, @Dasmat13)
  • KueueViz: Fix the workload detail page not showing the ClusterQueue name and preemption state. (#16905, @tenzen-y)
  • KueueViz: Fixed a bug where the workloads dashboard listed unrelated Pods for Workloads without a kueue.x-k8s.io/job-uid label, such as pod groups. Such Workloads now show an empty Pods list. (#16707, @henry3260)
  • Observability: Fix stale pending-workload metrics when LocalQueues are created, resumed, stopped, or deleted, including Workloads observed before their LocalQueue. (#16888, @tomsen02)
  • Observability: Fixed a bug where the kueue_replaced_workload_slices_total metric was not exposed on the metrics endpoint. (#16856, @Narwhal-fish)
  • Observability: Refresh pending-workload metrics when ClusterQueues stop or resume, without relying on subsequent Workload updates. (#16889, @tomsen02)
  • Ray integrations: Fixed a bug where workloads with worker groups using replicas and numOfHosts were not admitted because SubGroupCount and topology index labels were missing from the topology request. The KubeRayEvictOnInconsistentTopologyRequest feature gate is disabled by default, allowing existing workloads to continue running with old topology requests. (#15677, @spencer-p)
  • Ray integrations: Use zero instead of one as the worker replica fallback when replicas is absent, matching the KubeRay CRD default. (#16915, @mbobrovskyi)
  • Scheduling: Fixed a bug where Workloads blocked by a closed preemption gate could retain a stale QuotaReserved condition and status.resourceRequests. Such Workloads now report the PreemptionGated reason on the QuotaReserved condition instead of AdmissionGated. (#16761, @henry3260)
  • Scheduling: Fixed a bug where a CPU amount past the Quantity limit appeared in quota messages without the milli suffix, so it read as cores. (#16717, @HarshithaMS005)
  • Scheduling: Fixed a bug where a Workload whose total request exceeded the int64 range could be admitted against a quota it did not fit. (#16671, @HarshithaMS005)
  • Scheduling: Fixed a bug where partial admission of a PodSet whose total request overflowed int64 could admit more Pods than the quota allows, because the reduced request was derived from the saturated total instead of the per-Pod request. (#16520, @thc1006)
  • StatefulSet: Fixed a bug where changing a StatefulSet's queue while its Workload held quota could cause repeated Workload update failures when the StatefulSet had no ready replicas. Kueue now rejects the queue change until the Workload releases its quota. (#16710, @weizhoublue)
  • StatefulSet: Fixed a bug where labels listed in integrations.labelKeysToCopy, and the labels and annotations used as CustomMetricLabels sources, were not copied from a StatefulSet into its Workload. (#16770, @nsega)
  • TAS: Fixed a bug where a Workload using delayed topology assignment or failed node replacement could lose its quota reservation, and for a failed node replacement also its admission. (#16462, @apullo777)
  • TAS: Fixed a bug where a Workload using delayed topology assignment or failed node replacement could miss a scheduling retry until it was updated again. (#16468, @apullo777)
  • TAS: Fixed a bug where a second scheduling pass (failed-node replacement or delayed topology assignment) computed from an outdated Workload could overwrite a newer admission or drop status updated in the meantime, such as an Evicted condition. The scheduler now rejects the stale write, keeps the Workload's reservation, emits an OutdatedScheduleCycle event and retries the second pass. (#16859, @anguszzzz)
  • TAS: Fixed a bug where an elastic Workload with a TAS PodSet group could reassign its running leader to a different domain when the workers scaled up from zero, with the ElasticJobsViaWorkloadSlicesWithTAS feature gate enabled. (#16817, @Narwhal-fish)
  • TAS: Fixed a bug where stuck pods could break rank-based pod ungating after node replacement or workload re-admission. (#16795, @pajakd)
  • TAS: Fixed a scheduler panic when an elastic Workload scaled the workers of a PodSet group from 0 to the same count as the leader, with the ElasticJobsViaWorkloadSlicesWithTAS feature gate enabled. (#16812, @Narwhal-fish)
  • TAS: Fixed workloads with non-adjacent PodSet group members getting stuck during flavor-scan resume or delayed/failed-node topology reassignment. (#16658, @apullo777)

Other (Cleanup or Flake)

  • AgentSkills: Do not require the acknowledgment that the skills are experimental. (#16869, @mimowo)
  • ConfigAPI: Kueue logs configured internal labels in integrations.labelKeysToCopy at startup. Kueue does not copy these labels to a Workload. (#16878, @thc1006)

Don't miss a new kueue release

NewReleases is sending notifications on new releases.