What's New in v0.1.33
v0.1.33 adds per-container CPU and memory settings, a faster reaction to a usage surge or an OOMKill, and signed queries for Amazon Managed Prometheus. A namespace author can no longer use the operator's credentials to read another namespace or assume an AWS role you did not allow.
Highlights
- You can set CPU and memory per container, shorten the history window during a usage surge, raise memory after an OOMKill, and set a limit as a multiple of the request. Startup-boost samples can stay out of the CPU percentile. Memory HPA targets retune after a resize. Argo Rollouts are a target kind.
- Set the SigV4 allowlists before you upgrade if a policy or
AttuneNamespaceDefaultssetssigv4.roleArnorcloudwatch.roleArn. Both lists are empty until you set them.
Before you upgrade
Do this before the new controller starts. The checklist with the same steps is in Upgrading.
- Apply the v0.1.33 CRDs first. Helm does not update CRDs on
helm upgrade. Until those CRDs are in, a policy with nomaxAllowedfails its status write. - To keep the old ceiling, set
cpu.maxAllowedto4000mandmemory.maxAllowedto8Gion policies that omitted them. - A policy with its own
cloudwatchblock and nocpuUnitnow readscontainer_cpu_usage_totalas millicores. SetcpuUnit: Nanocoreson that policy to keep the old scale.AttuneDefaultsdoes not reach a policy that has its owncloudwatchblock. - Set
sigv4.allowedRoleArnswhen a policy orAttuneNamespaceDefaultssetsprometheus.sigv4.roleArnorcloudwatch.roleArn. Setsigv4.allowedWorkspaceHostswhensigv4has no role. ClusterAttuneDefaultsis not filtered. - A cross-namespace
metricsSource.vpa.namespacebecomesInvalidConfig. Move that VPA into the policy namespace, or put the cross-namespace read on clusterAttuneDefaults. - Update the controller ClusterRole with the chart or
dist/install.yaml.
New features
- Per-container CPU and memory. Each container can set its own percentile, bounds, and margin. A container you do not list keeps the policy defaults (#904).
- Usage surge window. A short window replaces the long history while usage is above the surge threshold, so a spike moves the recommendation sooner (#901).
- OOM bump. After an OOMKill on the original request, memory steps up by the configured bump instead of waiting for the next percentile (#899).
- Limit multiplier. A limit can be set as a multiple of the recommended request (#897).
- Startup history exclusion. Samples taken while a startup boost is in effect can be left out of the CPU percentile (#898).
- Memory HPA retune. After a memory resize, an HPA memory target is scaled from the new request the same way CPU already was (#903).
- Amazon Managed Prometheus.
prometheus.sigv4signs queries with the operator identity or a role ARN. A namespaced assume sendsExternalIdattune:<namespace>(#910). - Argo Rollouts.
target.kind: Rolloutresizes the live pods and can patch the Rollout template (#909).
Behavior changes
- An omitted
maxAllowedis no longer a hidden ceiling. The old defaults were4000mCPU and8Gimemory. A policy that does not setmaxAllowedis no longer capped there. Set those values yourself to keep the ceiling (#893). - CloudWatch CPU defaults to millicores. A policy
cloudwatchblock that omitscpuUnittreatscontainer_cpu_usage_totalas millicores. A throttle revert does not undo a bad CloudWatch step. SetcpuUnit: Nanocoresto keep the v0.1.32 scale (#892). - Namespace metrics are limited. A policy or
AttuneNamespaceDefaultsthat sets a SigV4 or CloudWatch role not on the allowlist becomesInvalidConfig. Tenant SLO guardrails that use operator Prometheus credentials are limited to the policy namespace. A query that already names another namespace is not sent. A scoped query with no samples does not revert, and theSLOGuardrailscondition reason isSLOGuardrailNoSamples. Datadog and CloudWatch stop evaluating guardrails. ClusterAttuneDefaultsis unchanged (#1033). - A rejected metrics source still unwinds work already done. A paused policy stays
Paused. An unpaused policy still reverts an OOM or a crash from a resize it already applied, and it still expires a startup boost that already landed. It does not query the rejected source or raise a new boost. A later revert is saved, so the same pod is not resized again every minute (#1034, #1035). - Native sidecars count in HPA resource sums. An init container with
restartPolicy: Alwaysis included. A stored HPA base that omitted that sidecar, and that still matches the old sum, is rewritten on the next successful retune (#1020, #1022). - The startup series cap keeps each container. When Prometheus returns more series than the cap, Attune keeps at least one series for every container instead of filling the cap with the first container's series. Policies outside
watchNamespacesare admitted (#1032). - Kubernetes 1.34 and newer can lower a memory limit in place. Older clusters still cannot. An in-place resize that would change the pod QoS class is skipped (#995, #894).
- A cooldown of
0sis rejected. A schedule window whose start equals its end is rejected. Reconcile does not stop on those objects (#888, #999).
Bug fixes
- Startup boost expiry lowered containers the boost had not raised. Expiry now changes only the containers listed on the boost stamp. A boost that lands in the same reconcile as another resize reads the live pod, so it does not apply twice (#1024, #1015).
- A canary could promote before a long SLO window finished. Promotion waits until that window has elapsed. An SLO
evaluationWindowlonger than the observation period is still evaluated (#1007, #1005). - An OOM at
maxAllowedwas counted again on every reconcile. It is counted once. An OOM bump hold no longer ends the safety observation (#1006, #1008). - A conflict could drop resize tracking or an HPA retune. Tracking is written with a merge patch. A failed HPA retune is retried. Observation stays open until the kubelet applies the resize (#1002, #1004, #1001).
- DaemonSet pods were left at the old size, and a rollout resize could land on pods that were about to be replaced. Current DaemonSet pods are resized. During a rollout, Attune still recommends and skips the resize only while pods are being replaced (#947, #895).
- CronJob pod names missed the minute stamp, and a memory limit could be a fraction of a byte. CronJob matching includes the minute. Scaled memory requests and limits round up to a whole byte (#987, #984, #1013).
- A stored
0son defaults disappeared on an unrelated update, and an inheritedoomBumprewrote defaults. The stored zero stays. Built-in defaults no longer copy over an inherited bump (#988, #911). - Datadog null points were stored as zero. Those points are dropped, so a gap does not look like idle usage (#887).
kubectl attune doctordid not name a missing Attune CRD. The CRD row says which object is missing (#1018).
Compatibility
| Surface | Requires |
|---|---|
Existing policies that omit maxAllowed
| No hidden ceiling after this upgrade. Set 4000m and 8Gi to keep the old cap
|
CloudWatch policies with their own cloudwatch block
| cpuUnit: Nanocores to keep the v0.1.32 scale
|
sigv4 or cloudwatch.roleArn on a namespaced object
| An allowlist entry before the new controller starts |
Cross-namespace metricsSource.vpa.namespace
| Invalid. Use the policy namespace, or cluster AttuneDefaults
|
| CRDs | Apply the v0.1.33 CRDs before the controller |
| Kubernetes | 1.32 through 1.36 required. 1.37 passed on 6dffe348 and is not required
|
| In-place memory limit decrease | Kubernetes 1.34 and newer |
Upgrade notes
- Upgrade the chart to 0.1.33, or set
image.tagto0.1.33orv0.1.33. - Pull
ghcr.io/attune-io/attune:v0.1.33orghcr.io/attune-io/attune:0.1.33. Both tags point at the same digest. - Apply the CRDs before the new controller. Helm does not upgrade CRDs on
helm upgrade. - Complete the steps in Before you upgrade.
See Upgrading for this release and earlier ones.
Contributors
Thanks to external contributors in this release:
- @konih for describing both budget cap events on the alert (#908).
- @konih for copying an inherited oomBump so built-in defaults do not rewrite it (#911).
- @konih for listing defaults in policy admission only when a surge window is set (#982).
- @konih for keeping the workloadref unread reason during template persistence (#983).
- @konih for rounding scaled memory limits up to a whole byte (#984, #985).
- @konih for matching CronJob pod names on the minute stamp (#987).
- @konih for keeping a stored 0s on defaults through unrelated updates (#988).
- @konih for reporting that a status write dropped an inherited cooldown, query step, and observation period (#986).
- @konih for reporting SLO and canary windows that ended too early (#972, #976).
- @konih for reporting an OOM at maxAllowed counted on every reconcile, and an OOM hold that ended observation (#974, #977).
- @konih for reporting tracking and HPA retune writes lost after a conflict (#973, #975).
- @konih for reporting the OperatorHub ClusterRole, unused list and watch verbs, cluster-wide Secret get, and the proxy-hop test gap (#978, #979, #980, #981).