github cloudposse/atmos v1.231.0-rc.3

pre-release5 hours ago
feat(scaffold): handle file deletions during scaffold/init --update @jorrite (#3245) ## what
  • atmos scaffold generate --update / atmos init --update under --update-strategy=rendered now deletes a file the template stopped generating between refs, but only when the on-disk copy is still byte-identical to the old pristine render. If local edits survived, the update fails with an unresolved merge conflict instead of silently deleting or keeping it.
  • Independently of --update-strategy, --update no longer silently recreates a file you deleted yourself — both tracked and rendered now check whether a path was previously generated before recreating it, and skip it if so.
  • New --recreate-deleted flag (on both scaffold generate and init) opts back into always recreating a deleted file. It's deliberately its own flag rather than folded into --force: --force already means "the template's version wins" for merge conflicts under --update, so tying recreation to it would make "manual conflict resolution" and "recreate what I deleted" mutually exclusive.
  • --force now also resolves a deletion conflict (deletes through local edits), consistent with that same existing meaning.
  • --update-strategy=tracked does not support template-side deletion propagation — documented as a limitation, not implemented, since there's no safe way to know which files in a target's git history actually belonged to the template.
  • Four fixes from a field-test pass on this change before it shipped:
    • a spec.files[].target: rename regression where the renamed file would never be created and the error message wrongly blamed the user for "deleting" it
    • a missing symlink guard on the deletion-candidate read path (now refuses to inspect a symlink, reusing ErrSymlinkWrite)
    • --force previously couldn't resolve a deletion conflict at all
    • the skip message now names --recreate-deleted as the way to override it
  • Unit tests (engine + UI layers) and real end-to-end tests driving the full CLI stack against scratch git-tagged template fixtures, covering both deletion directions, the conflict case, dry-run, and --recreate-deleted for both atmos scaffold generate and atmos init.
  • Docs: docs/prd/atmos-scaffold.md, website/docs/cli/commands/init.mdx, website/docs/cli/commands/scaffold/generate.mdx, plus a changelog post and roadmap entry.

why

  • A file the template stopped generating survived on disk forever, completely silently — the 3-way merge only ever considered files that existed on both the old and new side.
  • --update unconditionally overwrote a user's own deletion of a previously-generated file on every run, unlike tools like Copier that respect it.
  • A field-test pass (see commit history) caught a real regression this change would otherwise have introduced — breaking the existing, documented target: rename-migration path — along with a few consistency gaps, before any of it shipped.

references

  • No tracked GitHub issue; scoped directly from a task brief this session.

Summary by CodeRabbit

  • New Features
    • init --update and scaffold generate --update preserve files you deleted locally by default. Use --recreate-deleted to restore them with the template’s current content.
    • With the rendered update strategy, files removed from a template are deleted when unchanged. Locally edited files cause a conflict; --force deletes them.
  • Documentation
    • Updated command guides to explain file deletion and recreation behavior across update strategies.
feat(helm): surface crash-looping pod diagnostics on release failure (#3271) @aknysh (#3273) ## what
  • On a native-Helm release readiness failure, Atmos now enumerates the release's pods and folds their diagnostics into the same error it already builds: not-ready container status (CrashLoopBackOff/ImagePullBackOff, exit code, restart count), and - at debug/trace level - the failing container's log tail and the pod's recent events.
  • The diagnostics are captured before any rollback or uninstall deletes the pods. To guarantee that ordering, Atmos now performs the configured on_failure rollback/uninstall itself (after collecting diagnostics) instead of relying on Helm's inline RollbackOnFailure, which runs before Atmos sees the error.
  • Diagnostics are best-effort: any cluster-access error leaves the original failure unchanged. Verbose detail (log tail, events) is gated behind the log level, so normal output is unchanged.

Example (--logs-level=Debug):

Error: failed to perform helm release operation
  workload diagnostics:
    pod keda-operator-7d9f  keda-operator CrashLoopBackOff (exit 1, 5 restarts)
      last log (keda-operator):
        panic: failed to load config: invalid duration "5x"
      events:
        BackOff  Back-off restarting failed container

why

  • Closes #3271. The readiness-failure error (#2849) named the release but not the cause, so operators had to leave Atmos and run kubectl get pods/describe/logs - impossible in CI with no interactive cluster access.
  • When upgrade.on_failure: rollback (or install uninstall-on-failure) is configured, the recovery deleted the crashing pods first, destroying the evidence before anyone could look. Capturing diagnostics before recovery makes a failed controller, webhook, or gateway rollout legible from the Atmos output alone.

references

  • Closes #3271
  • Builds on the native Helm release lifecycle (#2849) and pairs with the server_side_apply/force_conflicts work
  • Fix write-up: docs/fixes/2026-10-04-native-helm-crashloop-diagnostics.md
  • Blog post: website/blog/2026-10-04-native-helm-crashloop-diagnostics.mdx; roadmap milestone under container-composition

Implementation

  • New pkg/component/helm/diagnostics.go: collectReleaseFailureDiagnostics lists pods by the standard app.kubernetes.io/instance=<release> label and summarizes not-ready init/regular containers. In verbose mode it also tails the failing container's log (the previous instance for a crash-looping container), bounded by both TailLines and LimitBytes, and the pod's recent events - narrowed server-side by field selector, sorted newest-first (with LastTimestamp/EventTime/CreationTimestamp fallbacks), and capped. Output is bounded throughout (max pods and events, message and log-byte caps); the clientset is behind a newReleaseClientset test seam.
  • pkg/component/helm/client.go: configureInstallLifecycle/configureUpgradeLifecycle no longer set Helm's inline RollbackOnFailure. On failure, installRelease/upgradeRelease collect diagnostics, then run Atmos-owned recovery:
    • rollbackFailedUpgrade rolls back to the latest successful (deployed/superseded) revision via lastSuccessfulRevision, mirroring Helm's own failRelease (plain version 0 could target an earlier failed revision), and honors the policy's chart_hooks/force_conflicts/server_side_apply; history is trimmed via enforceReleaseHistoryLimit.
    • uninstallFailedInstall honors chart_hooks, and install recovery runs only when a release record exists (releaseRecordExists) - a pre-mutation failure (bad chart, chart-load error, cancellation before apply) neither diagnoses nor uninstalls.
    • releaseOperationErrorWithDiagnostics folds the summary into the error. New ErrHelmReleaseRollback sentinel.

Testing

  • Collector unit tests via a fake clientset: crash-loop/image-pull/terminated/init-container summaries, verbose vs. non-verbose, release-label scoping, healthy/empty, clientset/pod-list/event-list errors, event newest-first ordering + bounding, and the event-timestamp fallbacks; plus the pure helpers.
  • Recovery tests: lastSuccessfulRevision (skips failed revisions; none-successful; no-release), releaseRecordExists, a pre-mutation install failure that skips diagnostics/uninstall, and end-to-end that a failed upgrade and install fold diagnostics into ErrHelmReleaseOperation before the Atmos-owned recovery (which still runs). Existing rollback/history/dry-run tests still pass.
  • go build ./..., go test ./pkg/component/helm/..., golangci-lint (0 issues), and cd website && npm run build pass.

Summary by CodeRabbit

  • New Features
    • Failed Helm installs and upgrades now include pod and container status diagnostics before configured rollback or uninstall.
    • Debug and trace output can include recent container logs and pod events. Diagnostics are best-effort and bounded, and successful operation output remains unchanged.
    • Configured upgrade rollback targets the latest successful revision. Dry-run failures skip cluster diagnostics.
  • Documentation
    • Added guidance and examples for Helm failure diagnostics and updated the product roadmap.

Don't miss a new atmos release

NewReleases is sending notifications on new releases.