Features
- dra prioritized alternatives support (#3142) #3142 (Ryan Mistretta)
- metrics: add karpenter_pods_disrupted_total counter (#3015) #3015 (praveen9354)
- metrics: add consolidation_policy and termination_mode labels to disruption metrics (#3179) #3179 (Derek Frank)
- add VPA prediction controller and store (#3134) #3134 (Jigisha Patil)
- default pod topology spread awareness (#3205) #3205 (Andrew Mitchell)
- account for instance type volume limits when provisioning (#3202) #3202 (Eddie)
- Reduce allocations of Requirement and Requirements (#3211) #3211 (Bevan Arps)
- metrics: document metric dimensions with built-in help and values (#3295) #3295 (Derek Frank)
- metrics: split the disruption decision dimension into command vs approval decisions (#3322) #3322 (Derek Frank)
- metrics: name and document the unknown cloudprovider error value (#3323) #3323 (Derek Frank)
- disruption: Add pod disruption tracking to perf test reporting (#2892) #2892 (Nathaniel Jones)
- node repair as voluntary disruption (#3311) #3311 (Derek Frank)
- terminate first drift for capacity constrained nodepools (#3312) #3312 (Derek Frank)
- publish pod event for insufficient capacity error (#3143) #3143 (Tsubasa Nagasawa)
- terminate-first repair for capacity-constrained NodePools (#3342) #3342 (Derek Frank)
- reboot as a node action (#3345) #3345 (Sarthak Umarani)
- Add feature-gated pod deletion cost management controller (#2894) #2894 (Nathaniel Jones)
- integrate reason-aware matching with voluntary node repair (#3336) #3336 (Sebestien)
- integrate node repair candidate resolution and escalation (#3367) #3367 (Sebestien)
- wire reboot into repair (#3377) #3377 (Sarthak Umarani)
- repair registered nodes that never initialize through the repair disruption method (#3393) #3393 (Derek Frank)
- add --legacy-node-repair to run the legacy node repair controller in place of the repair disruption method (#3400) #3400 (Derek Frank)
- add metric for terminate-first decisions labeled by why the replacement couldn't be staged first (#3404) #3404 (Derek Frank)
- always run the reboot controller and fail terminal reboot errors immediately (#3405) #3405 (Sarthak Umarani)
Bug Fixes
- fix capacity and overhead override semantic (#3151) #3151 (Jigisha Patil)
- heal stranded unlaunched nodeclaim entries in cluster state (#3190) #3190 (Ryan Mistretta)
- Filter out incompatible Offerings in EstimatedSavings() (#3217) #3217 (Drew Sirenko)
- propagate template labels to virtual buffer pods (#3174) #3174 (Jigisha Patil)
- check for nodename and terminal pod for PodSchedulingUndecidedTimeSeconds gauge delete (#3120) #3120 (Joshua Guo)
- relax minValues against narrowed NodeClaim requirements (#3170) #3170 (Karan V)
- adding dynamic resources field to nodeoverlay apply (#3278) #3278 (Joshua Guo)
- pass ctx to disruption.NewController in terminatefirst_test (#3346) #3346 (Sarthak Umarani)
- record terminate-first under the voluntary disruption decision label (#3348) #3348 (Derek Frank)
- handle not found error for node lookup in eviction queue (#3335) #3335 (Cameron McAvoy)
- guard
ReleaseNodeCountagainst a dropped NodePool entry (#3347) #3347 (GaneshBannur) - add nodepool label to cloud provider errors (#3269) #3269 (Cindia-blue)
- size repair replacements for pods whose eviction is blocked (#3358) #3358 (Derek Frank)
- add nodepool label to Reboot error metrics and expect unknown error label in tests (#3359) #3359 (Sarthak Umarani)
- don't record a reboot's terminal metrics twice on a stale reconcile (#3360) #3360 (Sarthak Umarani)
- capacity buffer: fix trigger when cb is deleted (#3298) #3298 (Zachary Nixon)
- disruption: emit DisruptionBlocked when static drift can't stage a replacement at the node limit (#3368) #3368 (Derek Frank)
- keep the reboot watch predicate from writing conditions into the cached NodeClaim (#3372) #3372 (Derek Frank)
- let residual pods ride a bounded reboot instead of deleting them (#3379) #3379 (Sarthak Umarani)
- capacity buffer: watch scalableRef workloads so buffers resolve as soon as their workload exists (#3370) #3370 (Derek Frank)
- correct go.tools.sum checksum for envoyproxy/go-control-plane/envoy v1.32.3 (#3388) #3388 (Sebestien)
- only repair uninitialized nodes once they register (#3397) #3397 (Derek Frank)
- terminate-first static repair and drift when capacity reservations are full (#3398) #3398 (Derek Frank)
- skip a static NodePool whose instance types can't be resolved instead of aborting repair and drift (#3403) #3403 (Derek Frank)
- bound the drain when a failed reboot escalates to replacement (#3407) #3407 (Sarthak Umarani)
- use the same timestamp to score all repair candidates (#3406) #3406 (GaneshBannur)
Documentation
- reference rfc-template.md in designs README (#3173) #3173 (Andrew Mitchell)
- add drift-per-nodepool-backoff rfc (#3128) #3128 (Cameron McAvoy)
- fix typos in contributing-guidelines.md (#3161) #3161 (Mridula Madabhushanam)
- Initial AGENTS.md (#3232) #3232 (Reed Schalo)
- RFC for default pod topology spread awareness (#3183) #3183 (Andrew Mitchell)
- add community maintained upcloud providers (#3309) #3309 (Leonardo)
- add AI disclosure section to PR template (#3334) #3334 (Drew Sirenko)
- RFC for making node repair voluntary disruption (#3192) #3192 (Derek Frank)
- RFC for Terminate-first Disruption for Capacity-Constrained Nodepools (#3203) #3203 (Derek Frank)
- RFC for reboot as a node action in karpenter (#3259) #3259 (Sarthak Umarani)
- RFC for Node Repair Veto (#3277) #3277 (GaneshBannur)
- RFC for building a pod deletion cost controller to enable Karpenter to work better with the ReplicaSet Controller (#2935) #2935 (Nathaniel Jones)
- add reason-aware repair policy matching RFC (#3263) #3263 (Sebestien)
- document well known annotations as wellknown.Annotations (#3364) #3364 (Derek Frank)
- add node repair candidate resolution, admission, and escalation RFC (#3270) #3270 (Sebestien)
- document feature gates as options.FeatureGates (#3376) #3376 (Derek Frank)
- remove consolidationPolicy default sentence from NodePool API docs (#3378) #3378 (Drew Sirenko)
Performance Improvements
- cache virtual pod definitions to prevent re-computation (#3137) #3137 (Zachary Nixon)
- rewrite Requirements.intersectKeys to avoid allocations (#3310) #3310 (Wiktor Kuropatwa)
- repair: match repair policies on Node updates instead of every disruption pass (#3381) #3381 (Derek Frank)
Tests
- add per-test-case performance threshold overrides (#3165) #3165 (Andrew Mitchell)
- add microbenchmarks for NewScheduler and NewTopology (#3299) #3299 (Ryan Mistretta)
- add catch for pod not found (#3315) #3315 (Ryan Mistretta)
- change foreground deletion to background (#3316) #3316 (Ryan Mistretta)
- add regression test for default pod tscs (#3274) #3274 (Andrew Mitchell)
- kwok: node repair + budgeted-breaker regression suite (#3357) #3357 (Derek Frank)
- disruption: cover terminate-first drift and repair (#3362) #3362 (Sebestien)
- e2e: patch repair fault conditions so stale node reads can't drop do not repair (#3371) #3371 (Derek Frank)
- add reboot e2e specs (#3380) #3380 (Sarthak Umarani)
- compare snapshotted container restarts across sequential runs (#3339) #3339 (Ryan Mistretta)
- perf: gate Karpenter memory on Go live heap instead of RSS (#3382) #3382 (Sebestien)
- raise reboot escalation e2e timeout to 35m (#3396) #3396 (Ryan Mistretta)
Chores
- complete operator enum validation for NodeSelectorRequirement (#2924) #2924 (Leo Ryu)
- default consolidateAfter instead of requiring it (#3220) #3220 (Ryan Mistretta)
- Consolidate pod force-deletion into the eviction queue (#3063) #3063 (Ketankumar Jani)
- upgrade KWOK to v0.8.0 (#3140) #3140 (Kotaro Inoue)
- clean up agents.md (#3264) #3264 (Todd Neal)
- controllers: standardize Name() to a string literal (#3268) #3268 (Derek Frank)
- use c.Name() for deviceallocation controller (#3283) #3283 (Derek Frank)
- improve controller error context (#3280) #3280 (Sebestien)
- breaking: standardize the reason metric label to lowercase values (#3284) #3284 (Derek Frank)
- Sync sig-autoscaling-leads aliases (#3293) #3293 (Daniel Kłobuszewski)
- breaking: set the queue_failures_total consolidation_type label to the command's consolidation type instead of its decision string (#3294) #3294 (Derek Frank)
- add dynamic_resources dimensions on pod metrics (#3286) #3286 (Ryan Mistretta)
- add stage field to metrics (#3321) #3321 (Ryan Mistretta)
- Add ryan-mist as reviewer (#3324) #3324 (Ryan Mistretta)
- add jigish620 as karpenter reviewer (#3325) #3325 (Jigisha Patil)
- push to PR cancels PRs inflight runs (#3349) #3349 (Ryan Mistretta)
- replace avast/retry-go with k8s wait utility (#2968) #2968 (Mikel Olasagasti Uranga)
- remove unused compatibility.karpenter.sh/provider annotation (#3375) #3375 (Derek Frank)
- don't register the VPA prediction controller until predictions are wired in (#3383) #3383 (Jigisha Patil)
- emit DisruptionBlocked when do-not-repair vetoes repair (#3390) #3390 (GaneshBannur)
- bump go to 1.26.9 and golang.org/x/net to v0.60.0 to fix vulncheck (#3392) #3392 (Derek Frank)