github attune-io/attune v0.1.33

3 hours ago

What's New in v0.1.33

v0.1.33 adds per-container CPU and memory settings, a faster reaction to a usage surge or an OOMKill, and signed queries for Amazon Managed Prometheus. A namespace author can no longer use the operator's credentials to read another namespace or assume an AWS role you did not allow.

Highlights

  • You can set CPU and memory per container, shorten the history window during a usage surge, raise memory after an OOMKill, and set a limit as a multiple of the request. Startup-boost samples can stay out of the CPU percentile. Memory HPA targets retune after a resize. Argo Rollouts are a target kind.
  • Set the SigV4 allowlists before you upgrade if a policy or AttuneNamespaceDefaults sets sigv4.roleArn or cloudwatch.roleArn. Both lists are empty until you set them.

Before you upgrade

Do this before the new controller starts. The checklist with the same steps is in Upgrading.

  1. Apply the v0.1.33 CRDs first. Helm does not update CRDs on helm upgrade. Until those CRDs are in, a policy with no maxAllowed fails its status write.
  2. To keep the old ceiling, set cpu.maxAllowed to 4000m and memory.maxAllowed to 8Gi on policies that omitted them.
  3. A policy with its own cloudwatch block and no cpuUnit now reads container_cpu_usage_total as millicores. Set cpuUnit: Nanocores on that policy to keep the old scale. AttuneDefaults does not reach a policy that has its own cloudwatch block.
  4. Set sigv4.allowedRoleArns when a policy or AttuneNamespaceDefaults sets prometheus.sigv4.roleArn or cloudwatch.roleArn. Set sigv4.allowedWorkspaceHosts when sigv4 has no role. Cluster AttuneDefaults is not filtered.
  5. A cross-namespace metricsSource.vpa.namespace becomes InvalidConfig. Move that VPA into the policy namespace, or put the cross-namespace read on cluster AttuneDefaults.
  6. Update the controller ClusterRole with the chart or dist/install.yaml.

New features

  • Per-container CPU and memory. Each container can set its own percentile, bounds, and margin. A container you do not list keeps the policy defaults (#904).
  • Usage surge window. A short window replaces the long history while usage is above the surge threshold, so a spike moves the recommendation sooner (#901).
  • OOM bump. After an OOMKill on the original request, memory steps up by the configured bump instead of waiting for the next percentile (#899).
  • Limit multiplier. A limit can be set as a multiple of the recommended request (#897).
  • Startup history exclusion. Samples taken while a startup boost is in effect can be left out of the CPU percentile (#898).
  • Memory HPA retune. After a memory resize, an HPA memory target is scaled from the new request the same way CPU already was (#903).
  • Amazon Managed Prometheus. prometheus.sigv4 signs queries with the operator identity or a role ARN. A namespaced assume sends ExternalId attune:<namespace> (#910).
  • Argo Rollouts. target.kind: Rollout resizes the live pods and can patch the Rollout template (#909).

Behavior changes

  • An omitted maxAllowed is no longer a hidden ceiling. The old defaults were 4000m CPU and 8Gi memory. A policy that does not set maxAllowed is no longer capped there. Set those values yourself to keep the ceiling (#893).
  • CloudWatch CPU defaults to millicores. A policy cloudwatch block that omits cpuUnit treats container_cpu_usage_total as millicores. A throttle revert does not undo a bad CloudWatch step. Set cpuUnit: Nanocores to keep the v0.1.32 scale (#892).
  • Namespace metrics are limited. A policy or AttuneNamespaceDefaults that sets a SigV4 or CloudWatch role not on the allowlist becomes InvalidConfig. Tenant SLO guardrails that use operator Prometheus credentials are limited to the policy namespace. A query that already names another namespace is not sent. A scoped query with no samples does not revert, and the SLOGuardrails condition reason is SLOGuardrailNoSamples. Datadog and CloudWatch stop evaluating guardrails. Cluster AttuneDefaults is unchanged (#1033).
  • A rejected metrics source still unwinds work already done. A paused policy stays Paused. An unpaused policy still reverts an OOM or a crash from a resize it already applied, and it still expires a startup boost that already landed. It does not query the rejected source or raise a new boost. A later revert is saved, so the same pod is not resized again every minute (#1034, #1035).
  • Native sidecars count in HPA resource sums. An init container with restartPolicy: Always is included. A stored HPA base that omitted that sidecar, and that still matches the old sum, is rewritten on the next successful retune (#1020, #1022).
  • The startup series cap keeps each container. When Prometheus returns more series than the cap, Attune keeps at least one series for every container instead of filling the cap with the first container's series. Policies outside watchNamespaces are admitted (#1032).
  • Kubernetes 1.34 and newer can lower a memory limit in place. Older clusters still cannot. An in-place resize that would change the pod QoS class is skipped (#995, #894).
  • A cooldown of 0s is rejected. A schedule window whose start equals its end is rejected. Reconcile does not stop on those objects (#888, #999).

Bug fixes

  • Startup boost expiry lowered containers the boost had not raised. Expiry now changes only the containers listed on the boost stamp. A boost that lands in the same reconcile as another resize reads the live pod, so it does not apply twice (#1024, #1015).
  • A canary could promote before a long SLO window finished. Promotion waits until that window has elapsed. An SLO evaluationWindow longer than the observation period is still evaluated (#1007, #1005).
  • An OOM at maxAllowed was counted again on every reconcile. It is counted once. An OOM bump hold no longer ends the safety observation (#1006, #1008).
  • A conflict could drop resize tracking or an HPA retune. Tracking is written with a merge patch. A failed HPA retune is retried. Observation stays open until the kubelet applies the resize (#1002, #1004, #1001).
  • DaemonSet pods were left at the old size, and a rollout resize could land on pods that were about to be replaced. Current DaemonSet pods are resized. During a rollout, Attune still recommends and skips the resize only while pods are being replaced (#947, #895).
  • CronJob pod names missed the minute stamp, and a memory limit could be a fraction of a byte. CronJob matching includes the minute. Scaled memory requests and limits round up to a whole byte (#987, #984, #1013).
  • A stored 0s on defaults disappeared on an unrelated update, and an inherited oomBump rewrote defaults. The stored zero stays. Built-in defaults no longer copy over an inherited bump (#988, #911).
  • Datadog null points were stored as zero. Those points are dropped, so a gap does not look like idle usage (#887).
  • kubectl attune doctor did not name a missing Attune CRD. The CRD row says which object is missing (#1018).

Compatibility

Surface Requires
Existing policies that omit maxAllowed No hidden ceiling after this upgrade. Set 4000m and 8Gi to keep the old cap
CloudWatch policies with their own cloudwatch block cpuUnit: Nanocores to keep the v0.1.32 scale
sigv4 or cloudwatch.roleArn on a namespaced object An allowlist entry before the new controller starts
Cross-namespace metricsSource.vpa.namespace Invalid. Use the policy namespace, or cluster AttuneDefaults
CRDs Apply the v0.1.33 CRDs before the controller
Kubernetes 1.32 through 1.36 required. 1.37 passed on 6dffe348 and is not required
In-place memory limit decrease Kubernetes 1.34 and newer

Upgrade notes

  1. Upgrade the chart to 0.1.33, or set image.tag to 0.1.33 or v0.1.33.
  2. Pull ghcr.io/attune-io/attune:v0.1.33 or ghcr.io/attune-io/attune:0.1.33. Both tags point at the same digest.
  3. Apply the CRDs before the new controller. Helm does not upgrade CRDs on helm upgrade.
  4. Complete the steps in Before you upgrade.

See Upgrading for this release and earlier ones.

Contributors

Thanks to external contributors in this release:

  • @konih for describing both budget cap events on the alert (#908).
  • @konih for copying an inherited oomBump so built-in defaults do not rewrite it (#911).
  • @konih for listing defaults in policy admission only when a surge window is set (#982).
  • @konih for keeping the workloadref unread reason during template persistence (#983).
  • @konih for rounding scaled memory limits up to a whole byte (#984, #985).
  • @konih for matching CronJob pod names on the minute stamp (#987).
  • @konih for keeping a stored 0s on defaults through unrelated updates (#988).
  • @konih for reporting that a status write dropped an inherited cooldown, query step, and observation period (#986).
  • @konih for reporting SLO and canary windows that ended too early (#972, #976).
  • @konih for reporting an OOM at maxAllowed counted on every reconcile, and an OOM hold that ended observation (#974, #977).
  • @konih for reporting tracking and HPA retune writes lost after a conflict (#973, #975).
  • @konih for reporting the OperatorHub ClusterRole, unused list and watch verbs, cluster-wide Secret get, and the proxy-hop test gap (#978, #979, #980, #981).

Full changelog

v0.1.32...v0.1.33

Don't miss a new attune release

NewReleases is sending notifications on new releases.