github longhorn/longhorn v1.13.0
Longhorn v1.13.0

4 hours ago

Longhorn v1.13.0 Release Notes

The Longhorn team is excited to announce the release of Longhorn v1.13.0. This feature release brings live upgrade to the V2 Data Engine, which reached general availability in v1.12.0. V2 volumes can now stay attached while Longhorn is upgraded, as long as your cluster meets the live upgrade prerequisites.

Longhorn v1.13.0 also introduces full interrupt mode for the V2 Data Engine, volume group snapshots, age-based retention for recurring jobs, topology-constrained replica scheduling, a scheduler extender, and the longhorn-global-manager Deployment that lowers control-plane overhead in large clusters.

For terminology and background on Longhorn releases, see Releases.

Breaking Changes

Kubernetes v1.34 Minimum Version

Because the CSI external-provisioner is upgraded to v6.3.0, all clusters must be running Kubernetes v1.34 or later before installing or upgrading to Longhorn v1.13.0.

Legacy V2 Linked-Clone Volumes

Since v1.12.1, V2 linked-clone volumes created in Longhorn v1.12.0 or earlier can only be detached or deleted. To replace one, create a new linked clone from the same source volume; no data copy is required.

GitHub Issue #12552

Primary Highlights

V2 Data Engine

The V2 Data Engine became generally available in Longhorn v1.12.0. Longhorn v1.13.0 adds live upgrade and full interrupt mode, and enables CPU isolation by default for new installations.

Important

V2 Volume Attach Latency at Scale:

In clusters with many attached V2 volumes, attaching additional volumes takes longer. The root cause is under investigation. For follow-up status, see Issue #13241.

ARM64 NVMe-backed Block-Type Node Disk Limitation:

On ARM64 systems, V2 volumes may experience stuck I/O when SPDK is configured with two or more CPU cores and node disks use the NVMe driver. The root cause is under investigation. As a workaround, use AIO-backed node disks instead of NVMe-backed node disks on ARM64 systems. For follow-up status, see Issue #13243.

UBLK Frontend Kernel Limitation:

The UBLK frontend for V2 Data Engine volumes remains experimental. On kernel v6.17, attaching a UBLK volume can cause a kernel panic. Do not use the UBLK frontend on this kernel version. For more information, see Issue #13509 and UBLK Frontend Support.

For the expected behavior differences between V1 and V2 volumes and the feature support matrix, see V1 and V2 Volume Feature Support.

GitHub Issue #6229

Live Upgrade

Longhorn v1.13.0 supports live upgrade of V2 volumes, so the V2 instance managers can be upgraded without detaching the volumes, preventing disruption during Longhorn upgrades. For more information, see V2 Data Engine Instance Manager Upgrade.

Important

Upgrade path:

V2 live upgrade to v1.13.0 is supported only from Longhorn v1.12.2. When upgrading from v1.12.0 or v1.12.1, or when the prerequisites are not met, detach all V2 volumes and ensure their replicas are stopped before upgrading.

GitHub Issue #9104

Full Interrupt Mode

In Longhorn v1.13.0, interrupt mode for the V2 Data Engine is now fully event-driven. SPDK reactors wait for I/O events instead of polling continuously, with only low-frequency background checks remaining, significantly reducing CPU usage for idle or low I/O workloads. This improves the hybrid implementation introduced in v1.10.0, which still incurred a minimal, constant CPU load when volumes were idle. As before, interrupt mode may introduce slightly higher I/O latency than polling mode under sustained high-throughput workloads.

Polling mode remains the default. To switch, set data-engine-interrupt-mode-enabled to {"v2":"true"} while no V2 volumes are attached. For more information, see Interrupt Mode Support.

GitHub Issue #11662

CPU Isolation Enabled by Default

CPU isolation keeps interrupts and other kernel background work off the CPU cores used by the V2 Data Engine. Longhorn v1.13.0 sets data-engine-cpu-isolation-enabled to {"v2":"true"} by default for new installations; clusters upgraded from v1.12.x keep their existing value and can opt in by setting it to {"v2":"true"}. CPU isolation applies only in polling mode; Longhorn skips it automatically when interrupt mode is enabled. For more information, see Data Engine CPU Isolation Enabled.

GitHub Issue #13724, GitHub Issue #13973

Fast Volume Cloning

Since v1.12.1, fast volume cloning for the V2 Data Engine has used a linked-clone architecture: a linked-clone volume shares data blocks with its source instead of copying them. One source volume can back multiple linked clones. Linked-clone volumes support most operations available to regular volumes, including snapshots, expansion, replica rebuilding, backup, and use as the source of nested linked clones. Starting with v1.13.0, V2 linked-clone volumes also support restore operations. For more information, see Volume Clone Support.

GitHub Issue #12552

Storage Sharding (Experimental)

Introduced in v1.12.1, storage sharding is an experimental V2 Data Engine feature. Instead of keeping a full copy of the volume on every replica, it splits written data into data and parity chunks and spreads them across nodes. A volume can then be larger than any single disk or node, and tolerate the same number of failures with less disk space.

Sharded volumes do not support backup and restore, volume cloning, backing images, disaster recovery (DR) volumes, or live migration. This feature is for evaluation and testing only and is not recommended for production. For more information, see Sharding with Erasure Coding.

GitHub Issue #1061

SPDK iobuf Pool Size Configuration

Introduced in v1.12.1, the data-engine-iobuf-large-pool-size and data-engine-iobuf-small-pool-size settings control the size of the SPDK buffer pools used by the V2 Data Engine. Larger pools help avoid running out of buffers under high-queue-depth workloads. Because the pools are sized when SPDK starts, changing either setting recreates V2 instance manager pods that have no running engines or replicas.

GitHub Issue #13322, GitHub Issue #13674

Snapshots and Backups

Volume Group Snapshot Support

Longhorn v1.13.0 makes it easier to protect related volumes by letting you snapshot a group of volumes with a single request from the Longhorn UI, kubectl, or Kubernetes VolumeGroupSnapshot objects through CSI. The CSI path, which is disabled by default, also supports group backups.

Note

Snapshot Consistency: Each member volume is snapshotted independently, so the group is not captured at a single point in time. Application-level consistency across the group is future work (Issue #2128).

For more information, see Create a Snapshot Group and Enable CSI Volume Group Snapshot Support.

GitHub Issue #13349

Age-Based Retention for Recurring Jobs

Longhorn v1.13.0 adds an age-based retention policy for snapshot, backup, and system backup recurring jobs, letting you retain data based on its age instead of the number of items. Set retentionPolicy to age-based and retainAge to a duration such as 720h; each run deletes snapshots or backups older than the specified age. The count-based policy remains the default, and existing recurring jobs are unchanged after the upgrade.

GitHub Issue #12060

Smarter Scheduling

Volume Topology Constraint

Longhorn v1.13.0 adds the volumeTopology StorageClass parameter (any, zonal, or regional) to keep a volume's replicas within the zone or region where it was provisioned, including during rebuilds and replica count changes, so replicas stay close to the workload. Previously, zone labels only spread replicas across zones; a rebuild could still place a replica in a different zone from the workload. For more information, see Topology-Aware Provisioning.

GitHub Issue #13493

Scheduler Extender

Longhorn v1.13.0 improves pod scheduling by allowing kube-scheduler to check actual Longhorn disk capacity and, when possible, place a restarted pod on the node that already holds all of its replicas. This helps avoid scheduling decisions based on stale CSIStorageCapacity data during bursts of pod creation and improves data locality, especially with best-effort data locality. The scheduler extender runs inside longhorn-manager and requires a kube-scheduler configuration change, which is not possible on managed Kubernetes services such as GKE and EKS.

GitHub Issue #12591

Better Operations

Longhorn Global Manager

Longhorn v1.13.0 reduces kube-apiserver load and longhorn-manager memory usage in large clusters by moving the cluster-wide pod and PersistentVolume controllers from the longhorn-manager DaemonSet to a new longhorn-global-manager Deployment. The Deployment is created with three replicas by default during installation and upgrade. Before upgrading, make sure at least one of its pods can be scheduled. For more information, see Upgrading Longhorn Manager.

GitHub Issue #13059

Security Hardening

Internal Network Policies

Since v1.12.1, Longhorn creates ingress NetworkPolicy resources for its internal components by default. They take effect only when the CNI plugin enforces NetworkPolicy. Longhorn v1.12.2 resolves the CNI compatibility issues found in v1.12.1 and adds three Helm values:

  • networkPolicies.v1DataEngineInitiatorSourceCIDRs and networkPolicies.recoveryBackendAdditionalIngressPorts allow traffic that the v1.12.1 policies blocked.
  • networkPolicies.metricsScrapeSources lets Prometheus and other scrapers reach longhorn-manager metrics on TCP port 9500.

For more information, see Internal Network Policies.

GitHub Issue #13438, GitHub Issue #13802, GitHub Issue #13740, GitHub Issue #13947

Instance Manager gRPC mTLS Coverage

Before v1.12.1, when the longhorn-grpc-tls secret was configured, mutual TLS (mTLS) covered only the instance manager's instance and proxy gRPC services; the disk and SPDK services still accepted plaintext connections. Since v1.12.1, mTLS covers all instance manager gRPC services, so every gRPC port requires a valid client certificate when the secret is configured.

GitHub Issue #7787, GitHub Issue #13212

Dedicated CSI Service Account

Longhorn v1.13.0 runs the CSI controller sidecars under a dedicated longhorn-csi-service-account instead of the shared longhorn-service-account. For compatibility with existing Secret references, the new service account still receives cluster-wide get access to Secrets by default. You can turn this off with the csi.allowControllerSecretAccess Helm value after removing csi.storage.k8s.io/provisioner-secret-* parameters from Longhorn StorageClasses. For more information, see Optional Restriction of CSI Controller Secret Access.

GitHub Issue #14020

Critical Stability Fixes

Linked-Clone Backup Restore

Longhorn v1.13.0 fixes a V2 linked-clone backup restore issue that could produce a corrupted volume when the source volume or the snapshot the clone was created from no longer existed. Backups of V2 linked-clone volumes now record the source volume and snapshot, and a restore fails with an error if either is missing.

GitHub Issue #13714

CSI Volume Clone with Strict-Local Data Locality

Longhorn v1.13.0 fixes a CSI volume clone issue where cloning a volume with dataLocality: strict-local could fail with hard affinity cannot be satisfied. Longhorn now picks a node that matches the volume's nodeSelector and diskSelector for the clone.

GitHub Issue #12792

Longhorn Node Removal After Kubernetes Node Deletion

Longhorn v1.13.0 fixes a node cleanup issue where a Longhorn node could not be removed after its Kubernetes node had been deleted without eviction. You can now disable scheduling on the node, remove its remaining replicas and engines, and then delete it. Evicting a node before deleting it from the cluster is still the recommended procedure. For more information, see Graceful Node Removal.

GitHub Issue #13494

Backing Image Copies on IPv6 Clusters

Longhorn v1.13.0 fixes a backing image issue where copies were always transferred over IPv4, so on IPv6 single-stack and IPv6-first dual-stack clusters a backing image stayed at one copy. Backing images are now copied over the cluster's IP family and the storage network.

GitHub Issue #13864

Installation

Important

Ensure that your cluster is running Kubernetes v1.34 or later before installing Longhorn v1.13.0.

You can install Longhorn using a variety of tools, including Rancher, kubectl, and Helm. For more information about installation methods and requirements, see Quick Installation in the Longhorn documentation.

Upgrade

Important

Ensure that your cluster is running Kubernetes v1.34 or later before upgrading from Longhorn v1.12.x to v1.13.0.

Longhorn only allows upgrades from supported versions. For more information about upgrade paths and procedures, see Upgrade in the Longhorn documentation.

Automated pre-upgrade checks do not cover all scenarios. Before upgrading, review the manual checks in Important Notes, and for V2 volumes confirm the live upgrade prerequisites or detach them and ensure their replicas are stopped before upgrading.

Post-Release Known Issues

For information about issues identified after this release, see Release-Known-Issues.

Resolved Issues in this release

Highlight

Feature

Improvement

Bug

Resilience

  • [BUG] V2 instance-manager liveness probe errors ('test: -eq: unary operator expected') -> self-kill -> node-wide replica fault under rebuild load 13957 - @yangchiu @mantissahz @Copilot
  • [DOC] Enhance Longhorn docs about Instance Manager 13197 - @Felipalds @chriscchien
  • [BUG] V2 Instance Manager panics in backupstore when S3 volume.cfg is missing, taking all node volumes offline 13790 - @mantissahz @chriscchien

Stability

Misc

Contributors

Don't miss a new longhorn release

NewReleases is sending notifications on new releases.