github LanternOps/breeze v0.117.0

4 hours ago

Breeze RMM v0.117.0 — Hardware & RAID monitoring, the alert/monitor conversion completes, a native Gmail ticket connector, and the billing-profile cutover finishes (breaking column drop).

⚠️ Breaking change: the six legacy labour-pricing columns (ticket_categories.default_billable/default_hourly_rate/rate_currency, org_ticket_settings.default_billable/default_hourly_rate/rate_currency) are dropped from the database in this release. See Self-Hosting / Upgrade Notes before upgrading.

Summary

Hardware & RAID monitoring ships end-to-end (#6854): the agent collects controller/array/disk health from storcli, mdadm, Windows Storage Spaces, smartctl and in-band BMC sources, the API rolls it up with dedicated alerting, and the device Hardware tab shows arrays, physical disks, cache batteries and collection status — attach the four built-in hardware monitors to a configuration policy to receive failure alerts. This feature needs the agent update in this release (0.117.0); devices on older agents show no hardware data. The Alerts → Monitors conversion (W05c) is now fully in the web UI: a Needs-conversion panel walks every legacy alert policy through an equivalence preview, Convert everything for partner managers, and a persistent Undo. Business Reports gains its W03 UI (#3198): SLA attainment, technician time & billability, and AR aging report types are now generatable from the Reports gallery, including a partner-wide "All organizations" option for partner managers. A native Gmail/Google Workspace inbound mailbox connector joins the existing Microsoft 365 one for email-to-ticket (#6740). Cove backup-provider integration reaches a working sync (#6008 W02/W03): third-party backup jobs now appear alongside first-party ones in backup health views. And the billing-profile cutover that began in v0.115.0/v0.116.0 finishes: the six legacy labour-pricing columns are archived and dropped (#4628 W04b, #6335) — the boot-time interlock refuses to start if any partner was never converted. The release also carries a wide fix sweep from the post-v0.116.0 pre-release sweep, plus a handful of fail-closed hardening fixes in AI tools and audit reads.

Added

  • Hardware & RAID monitoring (#6854): agent-side collection (storcli/perccli/megacli/ssacli/arcconf/omreport/mdadm/zfs/Windows Storage Spaces/smartctl, plus in-band BMC), API rollup and ingest, four built-in hardware monitors (RAID health, disk health, cache battery, BMC), a device Hardware tab (arrays, physical disks, controller cache batteries, collection status), a device-list column/filter, and a Hardware Monitoring section on configuration policies (collection on by default — RAID every 10 min, disk health every 60 min; monitors are not attached to a policy by default). Requires the updated agent (0.117.0); older agents report no hardware data.
  • Business Reports W03 (#6813, feature #3198): the three W02 report types — SLA attainment, technician time & billability, and AR aging — are now generatable from the Reports gallery (Business group), as PDF or CSV, on demand or scheduled. A partner manager (canManagePartnerWide) can generate a report across all organizations; other partner users must pick a single org.
  • Read-only report history for inactive organizations (#6771, #6850): an active partner can now read report definitions and past run metadata for its own suspended, churned, offboarding or archived orgs. Generating, scheduling, exporting or downloading stays refused.
  • Cove backup-provider integration, W02/W03 (#6008): a working sync job maps vendor customers/devices to Breeze orgs/devices, persists a daily health ledger with two-poll alert hysteresis, and surfaces third-party backup jobs in the Integrations Backup tab, the Backup overview, and the device Backup tab alongside first-party backup coverage.
  • Native Gmail / Google Workspace inbound mailbox connector (#6740): mirrors the existing Microsoft 365 email-to-ticket connector using the org's Google Workspace domain-wide delegation credential. Inbound only.
  • Partner API — read-only alerts feed (#6847): GET /api/v1/partner-api/alerts, behind a new opt-in alerts:read scope, lists alerts across every org a partner service principal can reach with lossless incremental sync. No acknowledge/resolve through this scope.
  • MCP: partner-wide scripts readable by org API keys, with a content pin (#6867): an org-scoped API key can now see its partner's partner-wide scripts through run_script/get_script_details/list_scripts (read-only visibility, no new write access). get_script_details returns contentSha256; run_script accepts an optional expectedContentSha256 and refuses a mismatch before dispatch.
  • MCP: opt-in unattended Tier 3 execution (#6841): MCP_UNATTENDED_TIER3_PRINCIPALS lets specific, explicitly-listed principals (api_key:<id> / oauth_client_user:<client_id>/<user id>) run Tier 3 MCP tools without interactive approval. Default is nobody; RBAC, the execute allowlist, and the Tier 3 execution ledger still apply.
  • System page (#6768): Settings → System now has a Connections tab (live status of configured integrations) and a Deprecations tab (what this version's deprecations/removals mean for this instance).
  • Metric anomaly episodes W04 + org/partner toggle (#6650, #6538): device anomaly panels are now episode-based, resolve requests can record episode-level feedback, and the ml.anomalies.enabled flag has a real Settings toggle at partner and org level (still off by default).
  • AI tools: manage_ticket_checklist (#6930): checklists become AI-writable; replace_unticked is guarded against a checklist with a waiting Operator step.
  • AI budgets: per-tool rate-limit multiplier (#6476, #6852): Settings → AI Usage can raise (never lower) the effective per-minute rate limit for a specific AI/MCP tool.
  • Configurable interactive AI approval timeout (#6475, #6829): 5–60 minutes, partner default with an org override.
  • Billing: org-level payment-terms override (#6229): an org can now override the partner's default invoice payment-terms days; NULL inherits.
  • Billing: partner identity falls back to company details (#6228): a partner whose Billing letterhead override (phone/website/address) is blank now has new invoices and quotes fill the gap from Settings → Company → Company details instead of rendering blank.
  • Invoices: presentation frozen at issue (#6227): an invoice's theme and page size are now snapshotted the moment it's issued, instead of re-reading the partner's live settings on every render.
  • Billing profiles W04a: work-type pickers on mobile + Outlook add-in (#4628): the mobile ticket timer and the Outlook time widget can pick a work type; the Billables CSV export gains work_type and included_minutes trailing columns.
  • Devices: installed helper (Breeze Assist) version is persisted and shown (#6751), and Devices now opens on the Agent segment, remembering the last choice (#5874).
  • Network devices can change site, with a real (non-no-op) profile site-change action (#6766).
  • Customer Portal Network Visibility, PR 2/3 of #5861 — per-asset network visibility endpoint and asset-level alert/ticket enrichment (active alert count, highest severity, open ticket count), behind an independent, fail-closed flag.
  • Scripts can raise alerts from per-script exit-code severity mapping (#6690).
  • Org bulk import gets a sample CSV and an import-guide link (#6051).
  • Settings catalogue (#6220, #6994): /settings is now a full, searchable catalogue of every settings screen, grouped by area; the sidebar Settings menu is trimmed to the most-used entries plus a new More settings link. One shared gate predicate drives both, so they can't disagree on what a role can see.
  • Mobile: navigate and defer across pending approvals (#6212, #6996): the approval takeover screen gets Prev/Next paging ("2 of 5") and a "Later" defer action when more than one approval is pending.
  • Custom fields are now selectable in the fleet-wide advanced device filter and as an opt-in device-list column (#6594, #6990) — previously only visible on the per-device detail page.
  • Quotes: editable per-quote footer line (#6648, #6970), separate from Terms & Conditions, blank = inherit the partner/brand kit footer.
  • Redis memory telemetry (#6452, #6968): used_memory/maxmemory/usage-ratio are now on /metrics, with a throttled warning once usage crosses 80% (tunable) — the first operator signal used to be a failed login or a stuck queue.
  • AI: large-tool-result capture decoupled from the hosted AI-workspace flag (#6732, #7014): self-hosters can now opt into paging back oversized tool results (read_artifact) via BREEZE_AI_ARTIFACT_CAPTURE_ENABLED without needing the hosted-only workspace lane — requires an artifact blob store to be configured; off by default.

Improved

  • Winget updates with no applicable upgrade, and superseded WUA updates, are now skipped rather than reported as failed (#6910).
  • Patch policy run-now jobs are enqueued only after the policy write commits, closing a race where a job could fan out against a not-yet-committed policy (#6632).
  • Quote emails use the proposal title and customer name with warmer default copy (#6814).
  • Webhook enable/disable toggle works reliably and legacy headers survive an edit (#6767).
  • Configuration Policies → Monitors: attach/detach now preserves a monitor's inheritance link (#6842).
  • Agent monitoring watches are cleared when no policy applies to a device, instead of lingering (#2949).
  • Backup verify / test-restore downloads are bounded and preserve partial counts on very large S3-compatible snapshots (tens of thousands of objects), instead of silently reporting 0 files ok 0 files failed after a 2-hour timeout (#6598).
  • The backup GC bucket listing is now streamed instead of materializing the whole bucket in memory — root cause of a September production OOM (#6834); a follow-up bounds three more GC memory/cost paths (reconcile full listing, local walk, unlimited-cap sweep) (#6843, #6984).
  • Failed agent latest-version lookups now use a short TTL (10s) instead of caching the failure (#6629).
  • POST /mssql/restore and POST /hyperv/restore now dispatch asynchronously instead of blocking the request for up to 10 minutes and reporting failure while the agent is still restoring (#6437, #6973); their terminal results are now persisted to restore_jobs like other restore types (#6974, #6991).
  • Backup restore surfaces (Overview, Restore Wizard, cleanup history, recovery tabs) format raw backend values consistently instead of leaking driver-internal representations (#6496, #6979); SnapshotBrowser selection now carries through into the Restore Wizard (#6456, #6980).
  • Portal: quote/invoice Terms & Conditions collapse into a native, printable <details> block (with a table of contents for longer terms) instead of a long inline wall of text, and the sign panel moves above it (#6575, #6995).
  • Devices whose update offer is being withheld (e.g. behind a version gate) now show a persisted reason and "since" date on the device Info tab and an "Updates withheld" badge in the fleet update view, instead of only a one-time log line (#6449, #6998).

Fixed

  • Breeze Assist (helper) bootstrap on hosted prod: helper installers are now registered by the local-mode binary sync (#6872, previously zero component='helper' rows meant no hosted device could bootstrap or upgrade Assist); the agent stops relaunching Assist when it isn't installed, and install retries are bounded with a withdrawn-offer path (#6872, #6921, #6927); a delivered helper response is now preferred over a later session close instead of being dropped (#6918, #6987).
  • Settings → System → Deprecations no longer 500s on a string-typed first_seen_at (#6833).
  • MIN() reboot-required timestamp is coerced to a Date instead of leaking a raw driver value (#6835).
  • Security scans: bulk Quarantine/Remove/Restore on a threat no longer writes a terminal status before the agent has actually acted — status now stays at its real pre-action value until the agent reports back, fixing a race where an offline device's threat could show "quarantined" without ever being quarantined (#6685). Threat status/severity badges and pluralization are now correctly localized.
  • Invoice line group headers (ticket subject + category) are now frozen at the moment an invoice is issued instead of re-reading tickets/ticket_categories live — renaming, recategorising, soft-deleting or org-moving a ticket no longer rewrites an invoice the customer already received; the authenticated customer portal (org-scoped RLS) can now show the category at all (#6674, #6955, #6975).
  • A verified backup snapshot's file index is now reset and re-hydrated after the device or snapshot org changes (move-org, org merge) instead of keeping stale tenancy provenance that permanently refused every external-reference download for that snapshot (#6488, #6988).
  • SSO: a passkey/WebAuthn login asserting amr: ["phr"] (phishing-resistant, e.g. Pocket ID) is now accepted as satisfying "Trust this provider's MFA," instead of forcing a redundant second factor (#6137, #7006).
  • A pending hosted partner under a forced-MFA role no longer loops between the inactive-account screen, /, and MFA setup (#6627, #6967).
  • Org merge no longer fails with no merge policy registered for '<table>' when an EE extension (e.g. Workspace) is enabled — extensions can now declare their own org-merge policy for their cascade-registered tables (#4165, #7005).
  • Live Linux bare-metal restore no longer copies the source machine's breeze-agent/breeze-watchdog systemd units and /etc/breeze onto an unenrolled recovery target, which previously left the target crash-looping trying to start services it has no agent installation for (#6436, #6999).
  • Windows agent: fixed a rare native crash from a recovered nil-pointer fault overflowing a small exception-frame stack reserve on AVX-512/AMX Intel hosts, hit only via a typed-nil error path added in #6883 (#6943, #6997).
  • Bare-metal recovery media now learns the server's minimum required ISO version before the one-time recovery code is spent, instead of burning the code on an outdated ISO and leaving the recovery stuck (#5629, #7013).
  • AI Operator: the five task-wide policy budgets (v15 spec) are now actually enforced at admission and dispatch instead of being computed but never checked (#6590, #7016).
  • MCP: a partner-scoped session can now see and call BYO-MCP org-owned tenant tools (previously silently omitted from tools/list and refused with Unknown tool on tools/call) (#6046, #6992); tools/list pagination now rejects a cursor from a stale tool catalog instead of silently skipping or repeating tools (#6407, #6986).
  • The PAM-reconciliation rate limiter no longer runs its Redis round-trip inside the request's held DB transaction (#6260, #6976).
  • A large post-v0.116.0 pre-release sweep also fixed: AI alert-triage duplicate/no-action notifications (#6750, #6963; #6449, #6932-linked run), monitor conversion equivalence re-validation on partial sourceIds (#6444, #6981), fleet-designer and device-detail paper cuts, billing stale-write guards on time-entry/ticket-part deletes, portal report-type filtering to exclude msp_staff (#6941, #6965), dev-push uploads staged under DEV_PUSH_WORK_DIR instead of /tmp (#6621, #6969), and numerous web admin surfaces (toasts, 404 modals, double-fetching alerts, autofill, locale-change re-renders, AI chat first-message visibility while streaming) — see PRs #6939–#6985, #7004, #7007, #7011, #7012 for the full list.

Security

  • audit_logs.details is now redacted on GET /ai/admin/security-events, matching the write-side redaction already applied to new rows; only rows written before the redaction existed, or by call sites that bypassed it, were affected (#6577).
  • AI org/site tools (list_organizations, list_sites, get_site, list_org_contacts, add_contact) and shared site-scope auth helpers now fail closed on a malformed (null/non-array) allowedSiteIds instead of either 500ing or silently treating it as unrestricted access (#6737, #6790).
  • AI write tools no longer guess an owner org from the first entry of a caller's accessible-orgs list — a device-page chat could previously create a resource in the wrong customer's org for a technician with several assigned orgs. manage_configuration_policy updates now reject an attempted owner-org change instead of silently dropping it (#6667, #6668).
  • #1105 system-context DB escalations are retired from two more readers now covered by a partner-wide RLS SELECT branch, so those reads run under the caller's own tenant context instead of a bypass (#5199, #7008).

Self-Hosting / Upgrade Notes

Upgrade: bump BREEZE_VERSION in .env, then docker compose pull api web portal && docker compose up -d (or pnpm install when running from source).

Database / migrations: 21 new migration files, all idempotent and auto-applying on boot via autoMigrate. None does a large-table rewrite against existing hot data. Two do a bounded, chunked backfill on write (not a blocking rewrite): the invoice-line ticket-subject/category snapshot backfills existing non-draft invoice lines in 5,000-row chunks (2026-10-31-100300), and the backup-snapshot-origin tenancy-reset migration re-checks and resets any complete file index whose recorded org/device no longer matches its snapshot (2026-10-31-100700) — both are one-time, counted (RAISE WARNING), and no larger than the affected row set. The rest add new tables/columns only (hardware health, Gmail mailbox connector, partner API alerts feed with a keyset index built CONCURRENTLY, report history read policies, invoice presentation snapshot, org invoice-terms override, device update-offer-withheld state) plus one column drop:

  • 2026-10-29-100300-drop-legacy-labour-pricing-columns.sql is the breaking one. It archives every row still carrying legacy labour pricing into a new table, legacy_labour_pricing_archive (full snapshot, not just the skipped rows — skip_reason names why the v0.115.0 conversion didn't carry a value into a billing profile), then drops the six legacy columns. The migration refuses to run — the API will not boot — if any partner was never converted to billing profiles, or if manual DDL left a table with only some of its three legacy columns. A normal upgrade path (v0.115.0 → v0.116.0 → v0.117.0) cannot produce either state, but a self-hoster who skipped straight from an older release should run the v0.115.0 conversion first. Any report, BI query, or integration reading ticket_categories.default_billable/default_hourly_rate/rate_currency or org_ticket_settings.default_billable/default_hourly_rate/rate_currency directly from the database will fail after this upgrade — repoint it at billing_profiles/billing_profile_rules/org_billing_profile_assignments, or at the values stamped on each time entry.
  • Rolling back to v0.116 after this upgrade needs a database restore, not just a BREEZE_VERSION revert — the columns the v0.116 image expects no longer exist, and its tenant-export policy still lists them, so org data export on a v0.116 image against an upgraded database will fail. Take a DB backup immediately before upgrading.

No new required environment variables. New optional env vars, all defaulted off/unset:

  • MCP_UNATTENDED_TIER3_PRINCIPALS (default unset = nobody) — comma-separated list of principals allowed to run Tier 3 MCP tools without interactive approval. RBAC, the execute allowlist, and the Tier 3 execution ledger still apply regardless.
  • REDIS_MEMORY_MONITOR_INTERVAL_MS / REDIS_MEMORY_WARN_RATIO (default 0.8) / REDIS_MEMORY_CAPTURE_THROTTLE_MS / REDIS_MEMORY_MONITOR_DISABLED — tuning/kill-switch knobs for the new Redis memory watchdog, which is on by default with sane defaults and only logs/emits metrics; it never changes Redis's noeviction policy.
  • BREEZE_AI_ARTIFACT_CAPTURE_ENABLED (default false) — self-host opt-in for large-tool-result capture/read_artifact. Setting it true without an artifact blob store configured (ARTIFACT_S3_BUCKET_<REGION>/S3_BUCKET + keys) makes the API refuse to boot, so treat it as required only if you opt in.

Behavior changes & flags:

  • Hardware & RAID monitoring (#6854) needs agent 0.117.0 — devices on an older agent show no hardware data until updated. Hardware monitors are not attached to any configuration policy by default; attach the four built-ins under Hardware Monitoring to start alerting.
  • alerts:read is a new opt-in partner-API scope; nothing changes for existing service principals until it's granted.
  • ml.anomalies.enabled now has a real UI toggle at partner and org level; it remains off by default.
  • Business Reports "All organizations" generation is limited to partner users who can manage partner-wide state (canManagePartnerWide); other partner users generate one org at a time.
  • Threat status in /security/scans no longer flips to a terminal state (quarantined/removed/allowed) at the moment an action is queued — it now reflects the device's real status until the agent reports the command result. A device offline when an action was queued may show its prior status longer than before, which is the correct/truthful state.
  • Settings navigation changed: several sidebar entries (Roles, SSO, Access Reviews, Enrollment Keys, Custom Fields, Variables, Saved Filters) moved off the sidebar into the new /settings catalogue; old URLs still work.
  • SSO: an identity provider asserting amr: ["phr"] now satisfies "Trust this provider's MFA" alongside amr: ["mfa"] — only relevant to partners who already opted into that trust setting.
  • POST /mssql/restore and POST /hyperv/restore now return once the restore is dispatched, not once it completes; poll restore_jobs/the existing restore UI for the terminal result instead of waiting on the HTTP response.

⚠️ Breaking changes: the legacy labour-pricing column drop above. No other breaking changes in this release.

Full Changelog: v0.116.0...v0.117.0

What's Changed

Full Changelog: v0.116.0...v0.117.0

Don't miss a new breeze release

NewReleases is sending notifications on new releases.