github PegaProx/project-pegaprox v1.3.0
PegaProx 1.3.0

4 hours ago

PegaProx 1.3.0 makes PegaProx itself highly available: up to four instances form a group with one leader and warm standbys. Next to that come automated Proxmox VE installations, bulk actions on guests, a set of new alert sources, a second security pass after 1.2.0 and a fix for a silent data loss in the ESXi live clone migration.

Please read the behaviour changes before updating. There are thirty-eight, several take something away from a setup that works today, and the update dialog lists them as well. The full list with explanations is on the docs: Upgrading to 1.3.

🪽 PegaProx High Availability (#625)

  • A second PegaProx can run as a warm standby of the first. It is paired once with a one-time code, pulls the shared configuration every 30 seconds, shows the clusters live and read only, and is promoted to leader by an admin when the leader is gone. The old leader steps aside as soon as it sees it was replaced, also when it comes back later.
  • A group holds up to four instances: one leader that runs all automation and holds the configuration, up to two members the leader makes active to serve users (consoles, shells and SPICE open there, every change goes to the leader), and standbys that wait to take over.
  • A standby hands a signed-in user's changes to the leader, which checks the account against its own users and names the standby in the audit log (on by default, switchable per instance).
  • Members talk over calls signed with Ed25519 that are useless anywhere else or later. A removed member is refused everywhere and goes passive. A sync never silently wipes what a member holds: whatever it would replace is kept as an encrypted copy an admin can download or dismiss.
  • Schedules run in the group's time zone. LDAP and OIDC sign-ins on a standby go by the synced account and write nothing there.
  • Everything lives in a new High Availability tab in the settings, with a banner for every user on a standby, in all three layouts and all nine languages. Failover is manual and confirmed; the groundwork for automatic failover is included but not enabled.
  • Docs: PegaProx High Availability

🛟 Node recovery

  • PegaProx's own node recovery moves a guest only when every volume it starts with is on shared storage, and reads a dead node's guest configs from the cluster file system on other nodes instead of guessing.
  • A node that returns during its recovery keeps what is left and no guest runs twice. A recovery that could not finish is listed in the cluster's HA settings.
  • A new self-fence agent (its own unit, pegaprox-fence-agent) decides by quorum first and asks PegaProx only where quorum cannot decide.
  • Two-node and forced-quorum clusters set up from now on are recovered only after a fence per node that was read back; existing setups keep their behaviour behind an explicit "unsafe two-node recovery" switch.

✨ New

  • Automated Proxmox VE installations: PegaProx serves the answer file to the Proxmox installer, each profile with its own fetch token and every run tracked from the first fetch to the installer's post-installation webhook. A six-step guided setup hashes the root password on the server, the answer file editor refuses what the installer would refuse, and it has a page of its own in all three layouts. Docs
  • Bulk actions on selected guests (start, shut down, stop, reboot, snapshot, tags) in every layout, each guest checked on its own and the result shown per guest. A node can start, stop or migrate all its guests at once.
  • Bulk migration runs on the server, one guest at a time, a few at once or all at once, follows each guest and can be cancelled, also across page reloads (#952).
  • New alert sources: failed Proxmox tasks such as a failed vzdump, Ceph health, failing or lagging replication, a ZFS pool that is not ONLINE or shows errors, snapshots older than a set age, guests that no backup job covers, and rolling update start, finish and reboot hand-over (#716). Each is raised once, with a resolved notice, and can be muted per rule or per object.
  • Rolling updates can move a node's templates with its guests (#763) and let negative affinity rules give way for the run, switched back on afterwards (#954), in scheduled runs and in a node's maintenance as well. ProxLB pin tags rank a drain's targets instead of blocking it, with switches for strict pins and an automatic move back (#811).
  • Containers from OCI images on PVE 9.1 and newer, from a small catalog or any image reference (a technology preview, as in Proxmox).
  • Check connection runs read-only probes against a cluster (API and fallback hosts, TLS fingerprint, the token's privileges, versions, clock, quorum, SSH) and names a fix for each failed item.
  • Directory mappings and virtiofs shares, container features after creation, PCI and USB passthrough through resource mappings (also on API-token clusters) and a VirtIO RNG, each checked before it reaches PVE.
  • The All Clusters overview lists the storage of every cluster, guests without backup, and exports the guest inventory. Node temperatures work without lm-sensors, ZFS pools show their device tree with error counts and the last scrub.
  • A metrics.view permission lets a Prometheus scrape run without an admin token (#818); the exporter also reports storage, replication state, backup age per guest and guest disk I/O.
  • The HTTP API is described in OpenAPI, generated from the routes and naming the permission each one demands, and can be browsed inside PegaProx from the user menu (#104, #693).
  • Also: favorites as a group at the top of the sidebar, search by MAC address, IP and notes, a broadcast banner for everyone or chosen tenants and roles, Cloud as a layout choice at first login, a System theme (#743), plugins limited to chosen clusters (#642), PEGAPROX_CONFIG_DIR and PEGAPROX_LOG_DIR (#826), a per-cluster switch that turns SSH off entirely (#941), PBS servers with an API token (#805), Site Recovery boot screenshots as test evidence, a refreshed cloud image catalog, a release Docker image for 32-bit ARM and snapshot .deb builds of the Testing branch (#973).

🔒 Security

A second and a third pass after 1.2.0, over the remaining findings of the penetration test and the daily code scans, including defects in 1.2.0's own fixes. Among them:

  • The PVE console ticket goes only to PegaProx's own console server, and an API token secret is never offered as an SSH password.
  • A cluster edit can no longer send its stored credential to a new address or inject an SSH option.
  • A delegate with admin.users could demote an administrator and replace their password in one request.
  • Read routes next to gated write routes handed out every tenant's custom roles, pool permissions, VM ACLs and the cluster security audit.
  • An API token with a lower role than its account got its privileges back through VM ACLs, list routes and the cluster gate.
  • A deleted VM's ACL went to the next guest on that VMID.
  • Accounts limited to pools or guests could change root@pam, write shared storage content and place guests.
  • Migration paths passed credentials where others could read them, left node files readable and could destroy more than they copied.
  • storage.create was never enforced, replacing a second factor needed no password, and host key pins could be lost under concurrent writes.
  • CI actions and images are pinned, static downloads are verified, and the release images verify the packages they install.

Every fix ships with a regression test and a counter-proof against the unpatched code.

⚠️ Behaviour changes

Thirty-eight, listed in the update dialog and explained on Upgrading to 1.3. The ones most likely to be noticed:

  • Adding storage needs storage.create; a deleted or unreadable custom role grants nothing instead of the viewer set.
  • Pool grants on directory groups take effect: a user whose LDAP or OIDC group holds a pool grant is now confined to that pool.
  • API tokens stay at their own role everywhere, and can no longer create tokens or manage 2FA.
  • Accounts limited to pools or VM ACLs no longer get whole-cluster data, no longer upload ISOs or create guests, and clone only onto nodes of their own guests.
  • Changing a cluster's host, port, fallback hosts or TLS verification needs the password or token secret again.
  • Node recovery leaves a guest with a local volume where it is (a local ISO included), and new two-node clusters need a read-back fence.
  • The ESXi live clone switchover now takes longer: it stops the source, carries over every block changed since the snapshot and only then starts the target, so the downtime grows with disk size (see Fixes, #1124).
  • The cloud image ids fedora-40, debian-11 and alpine-319 are gone.

🐛 Fixes

  • ESXi live clone migration lost every write made after its snapshot (#1124). In "Live Clone" and "Auto" mode the copy came from a snapshot, and nothing written on the source during the copy or the confirmation hold reached the target. The switchover now stops the source, folds the snapshot back, compares every block and copies what changed, and only then runs the post-copy steps and starts the target. If anything fails before the target runs, the source is powered back on.
  • The block delta of the pre-sync migration copied only the head of each changed block and still reported success. Changed blocks now arrive whole and are read back before the run continues.
  • VirtIO SCSI guests failed with INACCESSIBLE_BOOT_DEVICE after driver injection (#823): vioscsi gets its own bus type and the device ids it really has, and the drivers go into the control set Windows boots from.
  • An Entra sign-in failed with HTTP 500 when a group had no display name (#962).
  • Console behind a reverse proxy fixed end to end (#945, #955, #956); token-only clusters open consoles (#955); a TLS error from another connection no longer ends a healthy console (#713); paste types the right characters on non-US keymaps (#959); the terminal follows the bind address (#957) and no longer writes its script to /tmp (#958); the node shell no longer pans the page sideways (#727).
  • Updates: Check for updates refreshes apt first; rolling updates reboot once (#953), evacuate only to nodes allowed by HA affinity and storage (#647) and read the debsecan urgency correctly (#827).
  • Config drift no longer reports an offline node's objects as removed (#968), and a failed read is not a removal. Reports load a whole week for Last Week and show times in browser time (#963, #939).
  • Container migration with a target storage (#808); V2P disks on block storage as raw on Proxmox 9 (#816); the SMBIOS auto-configurator no longer rewrites running VMs (#943, #969); pool grants on LDAP and OIDC groups reach the user (#940); a slow answer no longer overwrites the next detail view (#828); several dialog and login buttons work again (#810); rolling updates on XCP-ng no longer stop at the first node; the container dialog lists templates again; answers on kept-alive connections no longer wait for a delayed ACK; the cluster health check reads all storage in one call (#946).

🧭 Known issues

  • Snapshot-Iterative (zero downtime) ESXi migration is experimental and can report success with missing or wrong data on the target. Do not use it for guests whose data matters; use Live Clone or Live Mirror instead. A fix is planned.

✅ Quality

The automated suite runs more than 9,800 tests on every change (authorization, integration, SSRF, crypto rotation, i18n, frontend invariants). The release was verified against live PVE 9.2 and PVE 9.1 clusters and an ESXi 8.0.3 host.

Thank you to everyone who filed, fixed, translated and sponsored along the way. 💚

Thanks in particular to @miketate1985, @gyptazy, @chico-gonzales, @grupoaxium, @stanthewizzard, @robertdahlem, @davlaw, @lordkill61, @daglimmer, @Frisch12, @avsdev-cw, @falschgeldkind, @shepart, @dschelte, @nvaert1986, @hugobugomugo, @maxilee, @brngates98, @si458 and @mdobprv, who filed or fixed the issues closed by this release.

💛 Sponsors

PegaProx is AGPL-3.0 and built in the open. Huge thanks to our Platinum sponsors who keep it moving:

💎 Platinum

Don't miss a new project-pegaprox release

NewReleases is sending notifications on new releases.