Changes since v0.19.6:
Actions Required Before Upgrading
(No, really, you MUST read this before you upgrade)
- Minor releases: Review the
.0release notes for each new minor version you cross; see:v0.18.0,v0.19.0. - Patch releases: Review the patch release notes leading up to this version, but only within this minor release line; see:
v0.19.1,v0.19.2,v0.19.3,v0.19.4,v0.19.5,v0.19.6.
Changes by Kind
Feature
- Controllers: Increased the maximum concurrent API updates per reconcile from 8 to 32. Controlled by the HighMaxParallelismWithinReconcile feature gate. (#16387, @mimowo)
Bug or Regression
- AFS: Fixed a bug where workloads could be scheduled in priority order instead of fair-sharing order when Admission Fair Sharing was enabled. (#16307, @abhayjoshi201)
- CLI: Fixed a bug where
kueuectl list pods --for pod/NAMEprintedNo resources foundinstead of listing the group members, when the Pod group name was stored in thekueue.x-k8s.io/pod-group-nameannotation rather than the label. (#16408, @tenzen-y) - CLI: Fixed a bug where
list pods --for pod/NAMEprintedNo resources foundfor a Pod that is not part of a pod group. The command now lists that Pod itself. (#16378, @tenzen-y) - Cohorts: Fixed a possible controller hang after an invalid Cohort hierarchy cycle update. ClusterQueues in a cyclic Cohort now surface a clear
CohortCycleDetectedcondition instead of hanging. (#16376, @tenzen-y) - Helm: Fixed ServiceMonitor metrics scraping to verify the cert-manager-issued certificate when enableCertManager is true and no custom tlsConfig is set. (#16318, @HsiuChuanHsu)
- KueueViz: Fix Workloads page crash when no workloads exist by adding the missing Alert import. (#16396, @Dasmat13)
- LeaderWorkerSet & StatefulSet: Fixed a bug where, during a rolling update, Pods at the current revision were ungated before the Workload was admitted, so they could run without quota. Kueue now ungates these Pods only when the Workload is admitted. (#16401, @henry3260)
- LeaderWorkerSet: Fixed a bug where a rolling update with a percentage
maxSurge(such as50%) had its surge groups' Workloads deleted while those groups were still running. The value was read as zero, so Kueue treated the update as having no surge. (#14405, @thc1006) - MultiKueue: Fixed a bug where evicted elastic workloads left stale remote objects on worker clusters, preventing re-admission from the manager-cluster specification. (#16178, @kevin85421)
- MultiKueue: Fixed a bug where the quota automation condition could retain a stale ObservedGeneration after a ClusterQueue spec update. Kueue now refreshes it after successful evaluation. (#16403, @mbobrovskyi)
- Observability: Fixed Admitted events reporting 0s when quota was reserved before admission. Events now show the actual time between quota reservation and admission. (#16412, @tenzen-y)
- Observability: Fixed kueue_cohort_subtree_admitted_active_workloads after Cohort hierarchy changes so it reflects admitted Workloads in the current subtree. (#16437, @tenzen-y)
- Pod: Fixed a bug where Pod groups could receive different ResourceFlavor assignments depending on client-provided role-hash values. PodSets are now ordered by scheduling shape rather than role-hash, making flavor assignment deterministic. Controlled by the
PodGroupSchedulingShapeOrderingfeature gate (Beta, enabled by default). (#14809, @amirialy) - Scheduling: Fixed a bug that could leave workloads pending when preemption target selection omitted
podsquota or used the spec count for a partially admitted workload-slice replacement. Kueue now sizes preemption targets using the assigned counts and resources. (#16200, @apullo777) - Scheduling: Fixed a bug where preemption for workloads with non-adjacent members of a kueue.x-k8s.io/podset-group-name group could use incorrect resource requests or flavors, causing a scheduler panic or leaving a Workload pending. Kueue now matches PodSets by name. (#16273, @apullo777)
- StatefulSet: Fixed a bug where changing the
kueue.x-k8s.io/queue-namelabel after the Pods were ungated but before any was Ready made the StatefulSet reconciler fail on every retry, and could keep a Pod scheduling-gated during a rolling update. Kueue now only updates the queue name on Pods that are still gated. (#16335, @henry3260) - TAS: Fixed a bug where a Node or TAS ResourceFlavor change requeued inadmissible Workloads in every ClusterQueue instead of only the ClusterQueues that use that flavor. (#16152, @sohankunkerkar)
Other (Cleanup or Flake)
- FairSharing: Improved preemption performance by up to 20% by precomputing lendable Cohort capacity instead of recalculating it for each preemption candidate. (#15955, @venuchitta)