Breeze RMM v0.121.0: one home for every AI model decision, a chat model picker, bring-your-own OpenAI-compatible models with tool use, automatic AI failover, AI usage chargeback to your clients, and backup destination credentials encrypted at rest.
⚠️ Roll forward only. Backup destination credentials are now encrypted at rest, and existing rows are encrypted on the first start. Rolling back to an earlier release after that start would send encrypted values to agents as storage credentials and break S3 backups. Take your usual database backup before upgrading (#7718).
⚠️ Self-hosters using MCP_LLM_PROVIDER=openai-compatible: action required. IS_HOSTED=false must now be set explicitly, the model must support tool calling to serve chat, and localhost endpoints are refused. Details under Self-Hosting / Upgrade Notes (#7771).
⚠️ Deprecated, removed in v0.122: /ai/provider and the legacy AI reviewer, Office and agent model settings fields. The AI model registry (/ai/models) replaces them. Upgrade to v0.121 before v0.122 rather than skipping it. Details under Deprecations (#7606).
Security — action required
Self-hosted operators on 0.120.0 and earlier should upgrade. A user who can configure a webhook, alert channel or image-from-URL field, or whoever runs the URL it points at, could repeatedly stop the API process, interrupting service for every tenant on the instance (GHSA-vjqq-fmr4-m576, High). A technician limited to selected organizations could read, edit or delete ticket comments in the partner's other organizations (GHSA-7xq8-mcq6-j438, Low).
Security hardening
This release includes security hardening across:
- backup destination credential storage;
- AI tool output;
- ticket comment and access review permissions;
- API key authentication;
- patch policy and privileged-access approval handling;
- outbound request handling;
- Quick Support session files and folder permissions;
- dependency updates and dependency audit checks.
Self-hosters are encouraged to upgrade. (#7716, #7723, #7729, #7751, #7782, #7718, #7687, #7706, #7591, #7614, #7662, #7721)
Summary
- AI Providers & Models settings (#7701, #7700). Partner Settings → AI Providers & Models is now the single place for AI connections, the models your techs may use, and the default model per AI feature. Org Settings → AI → Model defaults can tighten that per client. Every AI feature now picks its model, key and price from these settings, and every model call is metered once.
- Chat model picker (#7763, #7774). Techs pick the chat model and its options (effort, thinking) from the models you allow. Switching model mid-chat resumes the same conversation when it fits; otherwise Breeze offers a linked new chat that starts from a summary. AI agents pick their model the same way.
- Bring your own OpenAI-compatible model (#7771, #7801). Connect vLLM, LiteLLM, Ollama, OpenRouter, a hosted gateway or another OpenAI-compatible endpoint, discover or enter its models, verify them, and use a verified model on every AI feature, tools included.
- Automatic failover (#7775, #7789). When a model is unavailable, rate-limited, overloaded, out of credit or its key is rejected, AI calls move down an ordered backup list (up to 5). Moving between Breeze credits and your own API key happens only if you allow it. AI agent runs use a model per stage: triage, analysis, remediation.
- AI chargeback (#7759). Bill clients for the AI usage Breeze meters: set coverage (billable, included or not billed), a markup, or a per-model price list on a billing profile. Usage closes monthly into invoice lines, and a new AI usage by client report shows it. Nobody is billed by default, and usage from before this release is never billable.
- Model quality view (#7767). Settings → AI Usage gains a Quality view beside Spend, by model, feature or model family.
- Quick Support (#7687). Standard (non-admin) Windows users can start Quick Support, the session keeps its files in its own folder, and the user sees an on-screen "is viewing your screen" indicator while a technician is connected.
- Windows viewer updates install again (#7726). The in-app update of Breeze Viewer on Windows downloaded but never installed; it now installs, and a failed update says which stage failed.
Added
- AI Providers & Models tab (
#ai-provider): connections with a residency switch, models (enable, price, verify, refresh), and defaults by feature including AI agent stage roles, backup list and the cross-funding switch (#7701, #7775). - Org Settings → AI → Model defaults: tighten-only override; blank inherits and shows the inherited value (#7701, #7775).
- Settings → AI Usage breakdown by model, feature, technician and org, with refusal rate and failovers, plus the Quality view (#7701, #7767, #7775).
- Chat composer model menu and
POST /ai/sessions/:id/continuefor a linked continuation chat (#7763). - OpenAI-compatible connections with discovery (
GET {base_url}/models), manual model entry and a verification harness; new models land disabled, unpriced and unverified until you price and verify them (#7771). - AI chargeback: an AI usage section in the billing profile drawer (Settings → Billing → Rates),
ai_usageinvoice lines (non-taxable by default; editable on a draft), and the AI usage by client business report with PDF (#7759). - AI agent runs can end
blocked(model_unavailable/model_refused), notified once per agent per day; an org with no usable model skips the run instead of failing it (#7700). - Platform admins: a Prompt variants card on
/admin/ai-models(#7767). - Configurable agent enrollment rate limit:
AGENT_ENROLL_RATE_LIMIT(default 10 attempts) perAGENT_ENROLL_RATE_WINDOW_SECONDS(default 60), per source IP. Raise it for mass deployments behind one NAT address (#7724). - Windows backup captures junctions (for example a redirected
Musicfolder) as links and recreates them on restore, instead of skipping them. Volume mount points and other reparse points are still skipped and named in the job warning (#7757). - Antivirus detection recognises Emsisoft, Webroot, ThreatDown and WithSecure (including F-Secure) (#7634).
- Partner API ticket scopes (groundwork): service principals can be granted two new Partner API scopes,
tickets:readandtickets:write. Neither is in any default set. The ticket endpoints arrive in a later release, so granting the scopes has no effect yet. Comments an integration posts will be credited to the integration by name, not to a user, and never trigger an automatic AI reply; tickets can record their id and link in an external PSA or ITSM (newticket_external_refstable) (#7490). - New optional env vars
AI_PLATFORM_INFERENCE_GEO(usorglobal) (#7700) andAI_GATEWAY_HEADERS_TIMEOUT_SECONDS(#7801).
Improved
- One metering path for all AI features. Chat, topology, helper, script builder, Office, ticket drafts, the script reviewer, AI agents and enrichment are all priced from the same registry rates and recorded per call. Per-spend budget alerts cover every AI feature (#7700).
- Model refusals in chat show the category, alternative models and a docs link instead of an empty answer (#7700).
- Disconnecting an AI connection is a soft disconnect: the key is removed and its models disabled, and chats bound to it report it unavailable instead of silently moving elsewhere (#7700, #7771).
- Legacy model editors retired. The Office allowed-models list and the script-reviewer model fields now point to AI Providers & Models (#7701).
- AI chat diagnoses before it blames. Before naming a process, a vendor or the Breeze agent as the cause of load, the chat traces it to its owning process and service, works from a list of what the Breeze agent actually runs, and restates its evidence before proposing to kill a process or another destructive fix (#7764).
- AI agent runs are no longer offered tools the agent guardrail always refuses, which saves turns (#7758).
- Invoice payments are serialized per invoice. Pay links, portal payments, manual payments, accounting imports and voids now take a per-invoice reservation, laying the ground for autopay. Nothing new appears in the UI (#7777).
- Metric history storage. Metric rollup rows that landed in the catch-all partition are moved into proper monthly partitions by the daily maintenance job, so their space is returned and month creation no longer scans them (#7664).
- Patches: a per-device install supersedes an older failed job for that patch, and a patch that installed but awaits a restart shows Installed, reboot required instead of pending (#7727).
- Add Device reuses one installer enrollment key per site instead of creating a new one on every download or link (#7761).
- Network topology overview reads like a network map: each site network is one card, with its gateway above the LANs that route through it, and every device is a tile showing its type, name, address and whether it has an agent. Duplicate gateway/subnet boxes and
endpoint <uuid>labels are gone, and graph reads are faster (#7762). - Custom fields are listed alphabetically everywhere (#7660).
- Startup tasks (built-in monitor provisioning, settings-secret sealing, backup storage key history) retry with backoff after a transient failure instead of waiting for the next restart (#7725).
- Agent log level: a raised log-shipping level now survives an agent restart, reaches the desktop helper, and the command reports the level applied (#7590).
- Invoice PDF and email include a unit price column, matching the web preview and portal (#7638).
- Quote tax is resolved when the quote is sent, not when it was created, so a sent quote carries the org's current rate (#7754).
- Destructive buttons meet WCAG AA contrast in light and dark themes (#7720).
- Connected Apps explains when MCP OAuth is not enabled on the server (
MCP_OAUTH_ENABLED=true) instead of showing a bare 404 (#7722). - Viewer: connecting to a device whose agent runs as a service, from a viewer without WebRTC, now explains that the device needs WebRTC and what to use instead (#7755).
- Cloud backup storage format documented: S3, B2, Azure and GCS objects keep a
.gzname but are stored uncompressed; Local/NAS objects are gzip-compressed. Rename, don'tgunzip, a downloaded cloud object (#7715). - AI tools on Quick Support devices: screenshots, screen analysis, computer control and diagnose are refused for a Quick Support session (
409 SCREEN_ACCESS_UNAVAILABLE_IN_QUICK_SUPPORT); use the remote desktop session, where the user sees the indicator (#7687).
Fixed
- Ollama and slow local model servers work as OpenAI-compatible connections: tool calls through Ollama are no longer dropped, the first-response wait is longer on self-host and configurable, and choosing a model without tool calling for chat says to choose a tool-capable model instead of showing a bare error (#7801, #7793, #7794, #7795).
- Breeze Viewer in-app update on Windows now installs. The update package could not be unpacked, so viewers stayed on their installed version and showed "will retry on next launch" (#7726, #7681).
- Human API keys no longer stop working when their creator signs out. A password change or reset, invite acceptance, an admin status change or an MFA factor change still invalidates them (#7591, #7489).
- Built-in alert rules ("Patch job failures", "Reboot pending too long", the policy and config-compliance rules) can no longer be created twice. Existing duplicates are merged once at upgrade (#7750, #7650).
- Automation webhooks: a correctly signed webhook call no longer returns
404 Automation not found(#7589, #7363). - Webhook settings no longer go blank once a webhook has delivered or retrying deliveries, and Test confirms the delivery was queued (#7806).
- Quick Support can start for a standard (non-admin) Windows user (#7687, #7620).
- Quick Support no longer writes config, state or logs into the installed agent's folders; the installed agent re-secures its config folder if another account created it (#7687, #7629, #7645).
- Breeze Assist trusts certificates in the operating system's store, so self-hosted servers on a private CA connect (#7657, #7550).
- Binary sync: a failed S3 upload at boot is retried, and the previous release's binary is never served under the new checksum (#7661, #7574).
- Deleting a backup destination with job history returns a clear 409 naming what blocks it (disable the destination instead) rather than a 500 (#7636, #7622).
- Cancelling a queued bare-metal rebuild restore closes the recovery, so the device can be rebuilt again (#7611, #7512).
- Backup readiness score no longer drops while verifications are still running; a verification stuck over 30 minutes counts as failed (#7749, #7495).
- Vulnerabilities: Ready in the findings list now means Remediate will actually install a patch on that device (#7752, #7499).
- Linux encryption: a root filesystem on LVM over LUKS (the default Ubuntu/Debian encrypted install) reports as encrypted (#7579, #7478).
- Linux recovery media: a network or TLS failure, or a wrong server URL, no longer reads as "That code did not work" or spends a code attempt (#7658, #7649).
- Passwordless SSO accounts that already hold an MFA factor can add another passkey or TOTP (#7635, #7369).
- Ticket mailbox (M365): when the post-consent check fails because the application access policy isn't applied yet, the mailbox shows the policy steps and Re-test instead of looping back to Reconnect (#7612, #7569).
- Disk cleanup: Select all is capped at the 200 paths a cleanup run accepts, with a note explaining the limit, instead of failing on Delete (#7580, #7469).
- Run again on a device's script history resolves devices in that device's organization, not the org picked in the top bar (#7581, #7479).
- Partner billing settings are read-only, with a notice, for users without access to all organizations, instead of failing on Save (#7595, #7517).
- Org Viewer sign-in no longer shows a hydration error or partner-only 403s (#7737, #7498).
- Portal Backups: the "protected" count matches the table below it (#7742, #7505).
- Device identity edits no longer show a false "changed elsewhere" banner after your own save (#7632, #7388).
- Device mTLS certificates: devices that got their client certificate at enrollment, from admin provisioning or from quarantine approval were missing from Breeze's certificate history, so in enforce mode they could not renew or had their current certificate refused. Every issued certificate is now recorded, and the upgrade repairs existing devices once. Each superseded certificate is revoked at Cloudflare in the background over the following hour or so. A device whose certificate can't be recorded falls back to token auth (#7616, #7431, #7432).
- Hyper-V crash-consistent backup handles VMs that are Off, Saved or Paused. Off and Saved VMs are exported as they are and stay off; a Paused VM is left Saved, with a warning on the job; a running VM is restarted even if its export fails (#7630, #7623).
- Remote terminal: typing quickly or holding a key no longer ends the session. Input over the per-minute ceiling (6,000 messages or 8 MiB) is dropped with a warning and the session stays open (#7596, #7475).
- Deleting a site that still has devices, including removed ones, explains why it can't be deleted and how many devices are in the way, instead of a server error. Purge removed devices (Devices → show removed → Delete permanently) or move active ones, then delete the site (#7588, #7471).
- Organization merge no longer fails when a merged-away report run is cited as service-deliverable evidence; the evidence moves with the report run (#7613, #7443).
- Invoices without online payment: the public invoice page and the client portal no longer show a Pay button, the invoice PDF no longer prints a "Pay online" line, and invoice emails include "View & pay" only when payment is available (#7610, #7509).
- Xero / QuickBooks sync: when the provider rejects a customer or item as invalid (for example a name Xero won't accept), the mapping shows the provider's reason once and stops, instead of retrying five times and reporting each retry as an error (#7659, #7292).
- Accounting workbench: confirming a mapping while the background sync is already running no longer shows an error; the row updates to Synced on its own (#7656, #7386).
Known issues
- A chat turn that fails on a provider error before any output can end without an error shown (#7785).
- The Models card does not refresh when background discovery or verification finishes; reload the page (#7781).
- Breeze Assist on Windows does not run in RDP or other non-console sessions (#7542).
Self-Hosting / Upgrade Notes
Upgrade command. Check how your install pins its images: grep '^BREEZE_API_IMAGE_REF=' .env.
- Digest-pinned (value ends in
@sha256:…, everyguided-setup.shinstall from v0.112.0): fetch the currentguided-setup.shfrommain, then runbash guided-setup.sh --upgrade 0.121.0. It verifies the signed image inventory, updatesBREEZE_VERSIONand all four image digests, then pulls and restarts. - Tag-form (value ends in
:${BREEZE_VERSION}): setBREEZE_VERSION=0.121.0, then rundocker compose pull api web portal && docker compose up -d. - Do not edit only
BREEZE_VERSIONon a digest-pinned install. It re-pulls the images you already run. --upgradedoes not change yourdocker-compose.yml. To set any of this release's new optional variables, add their lines to theapiservice'senvironment:block (or take the v0.121.0docker-compose.yml); a value only in.envdoes not reach the API.- See the Upgrade Guide.
Breaking changes and required actions.
- Roll forward only: backup destination credentials are encrypted at rest (#7718).
- Access keys, secret keys, session tokens and the equivalent Azure, GCS and B2 credential fields in backup destinations are now stored encrypted with your application encryption key. Bucket, region, endpoint and prefix stay readable.
- On the first start, a background sweep encrypts existing destinations and logs one line:
[startup] Backup destination credentials sealed: <sealed>/<scanned> config(s) (<n> changed concurrently, <n> failed)
Re-running finds nothing to do. - After that line, do not roll back to an earlier release. Older code would hand the encrypted values to agents as storage credentials and S3 backups would fail. Fix forward instead.
- Key rotation:
scripts/re-encrypt-secrets.tsnow covers this column.
MCP_LLM_PROVIDER=openai-compatibleinstalls (#7771). The instance-wide OpenAI-compatible chat path is replaced by a connection in the new model registry.-
Set
IS_HOSTED=falseexplicitly. IfIS_HOSTEDis unset, empty or anything other than a recognised self-host value, the API refuses to start with:MCP_LLM_PROVIDER: MCP_LLM_PROVIDER=openai-compatible is for self-hosted Breeze (set IS_HOSTED explicitly to false). On hosted — or with IS_HOSTED unset/invalid — it is refused: it would create a connection for every partner. Add an OpenAI-compatible connection under Partner Settings → AI Providers & Models instead.The stock
docker-compose.ymlpassesIS_HOSTED: ${IS_HOSTED:-false}, so a Compose install getsfalseunless it overrides it. Installs without Compose (systemd, Kubernetes, bare Node) must set it. Note that the stock Compose file does not passMCP_LLM_*to the API; if you use this path you already carry your own mapping. -
What happens at boot. Each partner gets one read-only, env-managed connection and a priced, enabled model (prices from
MCP_LLM_PRICE_INPUT_PER_M_USD/MCP_LLM_PRICE_OUTPUT_PER_M_USD), verified once. Chat is pointed at it once, only where chat used the platform default. Restarts and multiple replicas don't create duplicates, and a later admin choice is kept. Partners created after boot get their connection within 10 minutes. -
The model must support tool calling to serve chat. A model that fails verification for tool use cannot serve chat; chat then says the model can't use tools and to choose a tool-capable model under AI Providers & Models (#7801).
-
localhostis refused. An env URL on a loopback, link-local or cloud-metadata address is refused; the API still starts, bootstraps nothing, and retries at 1, 5, 15 and 60 minutes, then every 10 minutes:[envOpenAiBootstrap] MCP_LLM_BASE_URL refused: That host is not reachable from Breeze (loopback, link-local and metadata addresses are never allowed). — no partner was bootstrappedhttp://localhost:11434(the usual Ollama URL) is therefore refused. Use the Docker service name or the host's LAN IP. Private-network addresses overhttp://are allowed on self-host. -
MCP_LLM_API_KEYis optional, but must be at least 8 characters when set. -
Removing the variables keeps the connection: it stays active and read-only, and the drawer says:
This connection was set up from the MCP_LLM_* environment variables, which are no longer set. It cannot be edited; you can disconnect it.Delete the
MCP_LLM_BASE_URLline rather than blanking it; an empty value fails config validation and the API will not start (unchanged from earlier releases). -
Model server sizing. Breeze's chat requests are large: allow at least 70k tokens of context (more for long chats).
-
First-response wait. The model gateway waits up to 120 s for a model server to start responding on self-host (30 s on hosted). Slow or CPU-only local servers (long prefill) may need to raise
AI_GATEWAY_HEADERS_TIMEOUT_SECONDS, up to 600. The value is clamped to 5–600 s; an invalid value is ignored with one warning. When the wait runs out, the request returns a 504 that names the setting, and the API logs:
[modelGateway] upstream sent no response headers within the <N> ms response-header deadline (…)(#7801) -
Ollama tool calling works (verified with qwen2.5 3b and 7b on Ollama 0.35.0). LiteLLM in front of llama.cpp
llama-server --jinjaalso works (#7801). -
Chat on this path is now recorded in AI usage like every other AI feature.
-
Deprecations (removed in v0.122). The AI model registry replaces each of these. In v0.121 they still work (the retired fields are already ignored); in v0.122 /ai/provider returns 404, and a request carrying a retired field is rejected with a 400 naming the field and its replacement, with nothing in it applied (#7606).
/api/v1/ai/provider(GET,PATCH,DELETE,POST /key,POST /endpoint). The web app no longer calls it. Use the AI model registry API:GET /api/v1/ai/models; connect a key withPOST /api/v1/ai/models/connections, rotate it withPOST /api/v1/ai/models/connections/:id/key, set an endpoint withPOST /api/v1/ai/models/connections/:id/endpoint, disconnect withDELETE /api/v1/ai/models/connections/:id. The single partner default model is replaced by per-feature defaults:PUT /api/v1/ai/models/assignments.reviewerModelonPUT /api/v1/ai/script-policyandPUT /api/v1/partner/ai/script-policy. Use thescript_reviewerfeature default (AI Providers & Models → Defaults by feature).allowedModelson the AI for Office organization policy (PUT /api/v1/client-ai/admin/orgs/:orgId/policy). Use theoffice_chatfeature's permitted models, or narrow them per organization withPUT /api/v1/ai/models/orgs/:orgId/assignments.modelon AI agents (POST /api/v1/ai/agents,PATCH /api/v1/ai/agents/:id,POST /api/v1/ai/agents/preview), which also carried the legacy agent model allowlist. SendofferingIdinstead: an enabled model fromGET /api/v1/ai/models;nullfollows the AI agents default.- Don't skip v0.121. The one-time move of legacy AI settings (script reviewer, AI for Office and AI agent model choices) into the registry runs in this release. An install that jumps straight to v0.122 starts each partner from the platform default and loses those choices.
Other behaviour changes.
- AI model settings moved. Model choices previously set in the Office policy editor, the script-reviewer model fields and the partner AI provider default are migrated once, on first start, into AI Providers & Models.
BREEZE_AI_SCRIPT_REVIEWER_MODELandWORKSPACE_CONTENT_LLM_MODELare no longer read at runtime; their values are migrated at that first start.ANTHROPIC_MODELremains the bootstrap default. - Residency switch. Turning on the residency requirement under AI Providers & Models disables every AI feature whose model can't guarantee it (the confirm lists them). BYO OpenAI-compatible models are never residency-eligible.
AI_PLATFORM_INFERENCE_GEO(new, optional):usorglobal. Any other value takes platform models offline, and the API warns at boot.- AI chargeback is off until you configure it. Every billing profile starts as not billed. A daily job (05:28 UTC) closes the previous UTC month per org. Chargeable ledger rows are kept at least 156 days regardless of
AI_INVOCATIONS_RETENTION_DAYS. - Self-hosted installs without a billing service never mark AI usage for credit debit, and turning billing on later cannot debit a backlog.
- Human API keys and sign-out (#7591). A sign-out, role change, email change or org merge no longer invalidates a human API key; a password change or reset, invite acceptance, an admin status change or an MFA factor change still does. A role downgrade or membership removal still takes effect on the key immediately. Keys that stopped working after a sign-out on v0.120 (#7489) are not revived by the upgrade: mint a replacement.
- Duplicate built-in alert rules are merged once at upgrade (#7750). The oldest copy is kept; if any copy was switched off, the merged rule is off. Open alerts from the removed copies move to the kept rule, and where two were open for the same device, the extra one is resolved as
source_retired. The API logs#7650: …warnings with the counts. - Metric history drain (#7664). On the first daily metric-rollup maintenance run after the upgrade, rows in the catch-all
metric_rollups_defaultpartition are moved into monthly partitions. Rows past your 5-minute retention are discarded, as before. Until the move finishes (seconds to minutes, by size), those older points are missing from charts. Look for[MetricRollupMaintenance] … drained metric_rollups_default: moved=N discarded=M …. - Quick Support: the viewing indicator always shows during a Quick Support session, regardless of the device's remote access indicator setting.
Database.
- 45 migrations (idempotent; applied automatically on boot via
autoMigrateunlessAUTO_MIGRATE=false). Expect[auto-migrate] Applied 45 migration(s). Some files are named earlier than v0.120's newest; they still apply, in order. - Mostly additive DDL: new tables for the AI registry cutover, AI chargeback (charges, runs, claims, per-model rates), autopay foundation tables (off by default), new columns on
ai_invocations,ai_sessions,ai_budget_reservations,ai_agent_runs,billing_profiles,llm_egress_events,usersandapi_keys, new enum values (report type, invoice line source, three antivirus providers), metric-rollup maintenance functions, a unique index on built-in alert rules, and policy updates. For the Partner API ticket groundwork (#7490): a newticket_external_refstable (RLS enforced), a nullable origin column onticket_comments, andtickets:read/tickets:writeadded to the allowed service-principal scopes. An index ontopology_node_bindingsspeeds up topology graph reads (#7762). - Three indexes are built with
CREATE INDEX CONCURRENTLYoutside a transaction, so writes continue: onai_invocations,ai_budget_reservationsanddevice_commands. A build waits for long-running transactions to finish. If theai_budget_reservationsbuild is interrupted, the next start refuses with an "INVALID index" message: runDROP INDEX CONCURRENTLY ai_budget_reservations_session_turn_idx;and restart the API. The other two have no such guard, so check them after the upgrade:If either isSELECT c.relname, i.indisvalid FROM pg_index i JOIN pg_class c ON c.oid = i.indexrelid WHERE c.relname IN ('ai_invocations_chargeable_idx', 'idx_device_commands_install_patches_device_created');
false, drop it withDROP INDEX CONCURRENTLY <name>;and re-run theCREATE INDEX CONCURRENTLYfrom its migration (2026-11-26-100150-ai-invocations-chargeable-idx.sqlor2026-11-19-101100-device-commands-install-patches-index.sql). - Brief locks: CHECK constraints on
ai_invocationsare added without a scan and validated in separate steps while reads and writes continue. A few constraint and index steps onai_sessions,ai_budget_reservationsandllm_egress_eventsscan the table while holding a write lock; they took well under a second at 300,000 rows per table in our test, so on very large tables expect a few seconds of paused AI chat writes during the first boot. New columns onusersandapi_keystake a momentary exclusive lock, and the alert-rule unique index briefly blocks alert rule writes. Adding theticket_commentscolumn and re-validating its two CHECK constraints holds an exclusive lock on that table for the scan (about 50 ms at 300,000 rows in our test), and thetopology_node_bindingsindex build blocks topology binding writes while it runs. No migration sets a lock timeout, so a long idle-in-transaction session on these tables would delay boot. - Device mTLS certificate repair (#7616). One migration records every device's issued client certificate in the certificate history. It holds only a read lock on
devices, so agents keep checking in, but it runs inside the boot: on a large fleet it can add seconds (tens of thousands of devices) to minutes to the first start before the API serves. It always logs one line, even when there is nothing to do:
WARNING device mTLS history reconcile: demoted <D> stale active rows, imported <I> issued certificates, skipped <S> devices sharing a certificate with another device
Demoted certificates are revoked at Cloudflare by the revocation sweep, 200 every 5 minutes. Devices counted as skipped share one certificate with another device and are left as they were. - First start after the upgrade also runs two background jobs that never block
/health: the backup credential encryption sweep above, and a one-time move of each partner's AI settings into the model registry ([startup] AI model registry cutover sweep: … N partner(s) cut over, 0 failed).
Environment variables.
-
Required: none.
IS_HOSTED=falsemust be explicit only onMCP_LLM_PROVIDER=openai-compatibleinstalls (above). -
New, optional:
AI_PLATFORM_INFERENCE_GEO(usorglobal).AI_GATEWAY_HEADERS_TIMEOUT_SECONDS(5–600; default 120 on self-host, 30 on hosted).AGENT_ENROLL_RATE_LIMIT(default 10) andAGENT_ENROLL_RATE_WINDOW_SECONDS(default 60). Blank keeps the defaults; a non-integer or non-positive value stops the API from starting.
All four are mapped in both Compose files of this release.
-
Deprecated:
BREEZE_AI_SCRIPT_REVIEWER_MODEL,WORKSPACE_CONTENT_LLM_MODEL(no longer read). -
Changed meaning:
MCP_LLM_*(above). -
deploy/docker-compose.prod.ymlnow passes 82 more variables to the API (#7756), the same lines the rootdocker-compose.ymlalready carries: for example the MCP OAuth, Mailgun, Twilio, SMTP, M365, Cloudflare and Redis memory monitor settings,SESSION_SECRET,JWT_SIGNING_KEYRINGandSENTRY_TRACES_SAMPLE_RATE. If you run that file, a value in your.envthat it used to ignore now takes effect. Nothing to do if your.envholds only what you meant to set; check it for stale entries before upgrading. Installs on the rootdocker-compose.yml(includingguided-setup.sh) are unaffected.
Agents. Agents and Breeze Assist update to 0.121.0 through your normal update channel. Agent changes: Windows Quick Support for standard users, private session files, config-folder checks and the viewing indicator (#7687); Windows junctions in backups (#7757); crash-consistent Hyper-V backup of Off, Saved and Paused VMs (#7630); persistent log-level overrides (#7590); LUKS-under-LVM encryption status (#7579); new antivirus providers (#7634); recovery media error messages (#7658); an OpenTelemetry Go dependency update (#7706). Breeze Assist now trusts the OS certificate store (#7657).
Viewer. Windows viewers on 0.118–0.120 whose in-app update kept failing pick up 0.121.0 in-app on their next launch; no reinstall is needed (#7726). If an update still fails, the banner now names the failed stage and the log path (%LOCALAPPDATA%\com.breeze.viewer\logs\updater.log); as a fallback, download breeze-viewer-windows.msi from this release and run it once by hand.
Feature flags. No feature flag changes. Prompt variants ship in a staged state, so no AI prompt changes at upgrade. Autopay ships switched off for every partner.
What's Changed
- docs(ai): model registry W11 plan — quality view + prompt profiles (#7609) by @ToddHebebrand in #7702
- docs(ai): model registry W08 plan — legacy removal (#7606) by @ToddHebebrand in #7707
- docs(ai): mark W07 deferred in the model registry plan index (#7605) by @ToddHebebrand in #7708
- chore(release): clear the next-release draft after v0.120.0 by @ToddHebebrand in #7711
- ci(security): unblock CI — npm-audit exceptions (node-forge GHSA-86w9) + bump agent OTel Go (GO-2026-6505) by @ToddHebebrand in #7706
- docs: v0.120.0 sweep — upgrade notes, restore integrity, consent protocol 2, AI model registry by @ToddHebebrand in #7712
- fix(tickets): ticket comment policies check the parent ticket's organization for read, edit and delete by @ToddHebebrand in #7716
- feat(ai): model registry W03 — resolveModel cutover, single billing path (#7601) by @ToddHebebrand in #7700
- fix(release): print notarytool's output when it exits non-zero (#7710) by @ToddHebebrand in #7714
- docs(backup): cloud backup objects keep the .gz name but are not compressed (#7621) by @ToddHebebrand in #7715
- feat(ai): model registry W04 — AI Providers & Models settings, org model defaults, AI usage by model (#7602) by @ToddHebebrand in #7701
- docs(portal): address spec review nits (#7450) by @fabicarvano in #7670
- fix(backup): store backup destination secrets encrypted by @ToddHebebrand in #7718
- chore(deps): bump tauri from 2.11.6 to 2.12.0 in /apps/helper/src-tauri by @dependabot[bot] in #7674
- chore(deps): bump tauri-build from 2.6.3 to 2.7.0 in /apps/viewer/src-tauri by @dependabot[bot] in #7675
- fix(access-reviews): review item policies follow the parent review's owner by @ToddHebebrand in #7723
- fix(ai): tool results omit stored key material by @ToddHebebrand in #7729
- fix(ai): webhook and monitor tool results show endpoints without stored paths or header values by @ToddHebebrand in #7751
- feat(ai): model registry W10 — AI chargeback (client pricing, monthly close, invoice lines, usage report) (#7608) by @ToddHebebrand in #7759
- feat(ai): model registry W05 — chat model picker, switching, agent policy picker (#7603) by @ToddHebebrand in #7763
- docs(testing): AI model registry W05 lab gate L2 results by @ToddHebebrand in #7770
- fix(agent): Quick Support for standard users, private session files, config folder checks and a viewing indicator by @ToddHebebrand in #7687
- feat(ai): model registry W11 — model quality view + prompt profile variants (#7609) by @ToddHebebrand in #7767
- fix(web): chat model picker survives the first message and session switches (#7768, #7769) by @ToddHebebrand in #7774
- docs(testing): W05 L2 scenario-1 re-check after #7774 by @ToddHebebrand in #7776
- fix(api): outbound requests handle empty-body and out-of-range status codes by @ToddHebebrand in #7782
- feat(ai): model registry W09 — failover walk + agent escalation roles (#7607) by @ToddHebebrand in #7775
- docs(testing): AI model registry W09 lab gates L1–L3 by @ToddHebebrand in #7788
- fix(ai): chat fails over on low credit; background CLI calls never label a failed turn (#7784, #7786) by @ToddHebebrand in #7789
- feat(ai): model registry W06 — BYO OpenAI-compatible connections via the model gateway (#7604) by @ToddHebebrand in #7771
- docs(testing): W09 L1 low-credit re-check after #7789 by @ToddHebebrand in #7791
- docs(testing): AI model registry W06 lab gates L1–L3 by @ToddHebebrand in #7796
- docs(billing): autopay design spec + wave plans (#7743) by @ToddHebebrand in #7753
- docs(testing): AI model registry lab docs — status banners, corrections, redactions by @ToddHebebrand in #7799
- feat(autopay): W01 foundation — tables, tenancy, reservation, settlement, outbox (#7744) by @ToddHebebrand in #7777
- fix(viewer): Windows auto-update never installs — stored update zip + failed-stage diagnostics (#7681) by @ToddHebebrand in #7726
- chore(web): What's New entry for v0.121.0 by @ToddHebebrand in #7802
- fix(web): no false 'changed elsewhere' banner after own identity save (#7388) by @ToddHebebrand in #7632
- fix(patches): device patches route reads partner-axis rows on one connection (#7647) by @ToddHebebrand in #7662
- fix(agent/recovery): stop blaming the code for network/TLS failures (#7649) by @ToddHebebrand in #7658
- fix(api): never serve a stale S3 binary after a failed boot upload (#7574) by @ToddHebebrand in #7661
- fix(helper): trust the OS certificate store in Breeze Assist (#7550) by @ToddHebebrand in #7657
- fix(custom-fields): list definitions alphabetically (#7572) by @ToddHebebrand in #7660
- fix(metric-rollups): drain metric_rollups_default instead of row-deleting it (#7541) by @ToddHebebrand in #7664
- feat(api): configurable agent enrollment rate limit (#7472) by @ToddHebebrand in #7724
- fix(patches): per-device installs supersede stale failures; show installed-awaiting-reboot (#7680) by @ToddHebebrand in #7727
- fix(agent): report LUKS-under-LVM roots as encrypted on Linux (#7478) by @ToddHebebrand in #7579
- fix(agent): persist set_log_level overrides, reach helpers, report the applied level (#7416) by @ToddHebebrand in #7590
- fix(auth): bind human API keys to a credential epoch so logout no longer kills them (#7489) by @ToddHebebrand in #7591
- fix(api): run the automation webhook under a DB context — lookup as system, writes as the owner (#7363) by @ToddHebebrand in #7589
- fix(pam): one lock order for elevation approvals (#7526) by @ToddHebebrand in #7614
- fix(tickets): keep M365 mailbox in error when consent-callback probe fails for non-auth reasons (#7569) by @ToddHebebrand in #7612
- fix(backup): restore cancel closes a queued rebuild's bare-metal recovery (#7512) by @ToddHebebrand in #7611
- fix(web): partner billing settings read-only without partner-wide access (#7517) by @ToddHebebrand in #7595
- fix(api): 409 instead of 500 when deleting a backup destination with history (#7622) by @ToddHebebrand in #7636
- fix(billing): add unit price column to invoice PDF and email (#7508) by @ToddHebebrand in #7638
- fix(agent,api): recognise Emsisoft, Webroot, ThreatDown, WithSecure as AV providers (#7551) by @ToddHebebrand in #7634
- fix(web): Connected Apps explains when MCP OAuth is disabled (#7697) by @ToddHebebrand in #7722
- fix(security): npm audit gate fails closed on unrankable advisories and vacuous scans (#7704, #7705) by @ToddHebebrand in #7721
- fix(api): retry detached startup tasks with backoff (#7693) by @ToddHebebrand in #7725
- fix(web): destructive outline buttons meet WCAG AA contrast in both themes (#7511) by @ToddHebebrand in #7720
- fix(vulnerabilities): list Ready state uses Remediate's patch resolver (#7499) by @ToddHebebrand in #7752
- fix(backup): exclude in-flight verifications from readiness score (#7495) by @ToddHebebrand in #7749
- fix(quotes): resolve the quote tax rate at send, not at create (#7507) by @ToddHebebrand in #7754
- fix(ai): full-profile exposure drops tools the agent guardrail always refuses (#7447) by @ToddHebebrand in #7758
- fix(enrollment): Add Device reuses one parent key per site instead of minting one per click (#7345) by @ToddHebebrand in #7761
- fix(viewer): explain service-mode WebSocket refusal when WebRTC is unavailable (#7415) by @ToddHebebrand in #7755
- fix(compose): map 82 operator env vars into prod api container (#7640) by @ToddHebebrand in #7756
- feat(agent/backup): back up Windows junctions as links and recreate them on restore (#7325) by @ToddHebebrand in #7757
- fix(ai): gateway works with Ollama tool calls, slow local servers, and explains tool-less models (#7795, #7794, #7793) by @ToddHebebrand in #7801
- fix(auth): let a passwordless SSO account holding a factor add another (#7369) by @ToddHebebrand in #7635
- fix(web): Run again resolves devices in the viewed device's org (#7479) by @ToddHebebrand in #7581
- fix(web): cap disk-cleanup selection at the API's 200-path limit (#7469) by @ToddHebebrand in #7580
- fix(web): Org Viewer hydration error + partner-only 403s (#7498) by @ToddHebebrand in #7737
- fix(portal): one meaning of 'protected' — Backups count matches its table (#7505) by @ToddHebebrand in #7742
- fix(alerts): built-in anchor rules can no longer be created twice on concurrent first fire (#7650) by @ToddHebebrand in #7750
- fix(ai): attribute load before blaming, and ground the chat in what the agent really runs (#7582) by @ToddHebebrand in #7764
- docs(testing): v0.120.0 → main pre-release sweep + webhook delivery history fix by @ToddHebebrand in #7806
- docs(specs): Partner API tickets surface design (#7181) by @lennonflima in #7246
- feat(api): partner API ticket scopes, typed ticket actor, external refs (wave 1) by @lennonflima in #7490
- fix(accounting): post-decision sync_in_progress is queued, not a failure (#7386) by @ToddHebebrand in #7656
- fix(billing): hide Pay CTA and PDF pay-online line when online payment isn't set up (#7509) by @ToddHebebrand in #7610
- fix(accounting): make a code-less provider validation error terminal (#7292) by @ToddHebebrand in #7659
- fix(mtls): record every issued certificate in device_mtls_certificates (#7431, #7432) by @ToddHebebrand in #7616
- fix(org-merge): re-home deliverable evidence with merged report runs (#7443) by @ToddHebebrand in #7613
- fix(sites): 409 instead of 500 when deleting a site that still has devices (#7471) by @ToddHebebrand in #7588
- fix(terminal): drop excess input instead of closing session on rate limit (#7475) by @ToddHebebrand in #7596
- fix(agent): crash-consistent Hyper-V backup handles Off/Saved/Paused VMs (#7623) by @ToddHebebrand in #7630
- test(accounting): pin the clock for QuickBooks CDC fixtures (main is red) by @ToddHebebrand in #7832
- ci(security): unblock Trivy — node-forge CVE-2026-85393 alias + pip-vendored urllib3 by @ToddHebebrand in #7813
- fix(topology): grouped overview — network cards, live labels, device tiles by @ToddHebebrand in #7762
- fix(helper): match @tauri-apps/api to the tauri 2.12 crate by @ToddHebebrand in #7844
New Contributors
- @lennonflima made their first contribution in #7246
Full Changelog: v0.120.0...v0.121.0