IncidentRelay 2.3.1 introduces Alert Analytics, Prometheus observability, improved UI navigation, flexible configuration management, and expanded deployment support.
This release builds on Incident Management v2 introduced in 2.3.0, focusing on operational visibility, usability, reliability, and easier deployment in production environments.
Highlights
Alert Analytics Dashboard
Introduced Alert Analytics v1, accessible through the new Analytics tab on the Alerts page.
The dashboard provides:
- AlertGroup lifecycle trends showing created, acknowledged, and resolved alerts.
- Noisiest alerts, with direct links to matching alerts in the inbox.
- Oldest unresolved AlertGroups.
- Acknowledgement health and currently unacknowledged alerts.
- Analysis of alerts resolved without acknowledgement.
- MTTA and MTTR metrics, including p50 and p95 percentiles.
- Selectable 7, 30, and 90-day reporting windows.
- Team-scoped analytics respecting RBAC and global team filtering.
Suppressed AlertGroups are excluded from response-quality calculations to avoid misleading metrics.
A new analytics API, OpenAPI documentation, translations for all six supported languages, and demo-data generators are also included.
Prometheus Metrics and Observability
Added an optional Prometheus-compatible /metrics endpoint for monitoring IncidentRelay in production.
Available metrics include:
- HTTP request rates and response latency.
- Database availability and pending migrations.
- Alert intake volume and AlertGroup lifecycle transitions.
- Notification delivery statistics and recent errors.
- Scheduler, Telegram, and Slack worker heartbeats.
- User notification queue depth and backlog age.
- Event Orchestration pending, activating, and failed event counts.
- Age of overdue orchestration events.
- Application build and version information.
Additional improvements:
- Support for Prometheus multiprocess metrics collection.
- Database-backed worker heartbeats for cross-process visibility.
- Improved readiness checks and metrics availability during database outages.
- Efficient 24-hour notification metrics backed by database indexes.
- Improved handling of stale worker metrics and process restarts.
Metrics are disabled by default and can be protected using bearer-token authentication.
Configuration Through Environment Variables and Secrets
Added support for overriding configuration options using environment variables:
INCIDENTRELAY__DATABASE__PASSWORD=...
INCIDENTRELAY__LOGGING__LEVEL=DEBUGConfiguration values can also be loaded directly from files:
INCIDENTRELAY__DATABASE__PASSWORD__FILE=/run/secrets/db-passwordThis enables easier integration with Docker Secrets, Kubernetes Secrets, and external secret-management systems.
Improvements include:
- Environment variables take precedence over configuration-file values.
- New Helm
configFromsupport for Secrets, ConfigMaps, and mounted files. - Consistent configuration delivery across web, scheduler, and worker components.
- Improved handling of shared application security keys.
- Docker runtime configuration is refreshed on restart, allowing configuration changes to take effect without recreating persistent configuration files.
Existing file-based configurations continue to work without modification.
Docker and Kubernetes Improvements
Multi-architecture Docker images
Official Docker images now support:
linux/amd64linux/arm64
Images are published as multi-architecture manifests, improving support for ARM64 servers and Kubernetes clusters.
Helm ServiceMonitor integration
Added an optional Prometheus Operator ServiceMonitor to the Helm chart.
Features include:
- Configurable scrape interval, timeout, and labels.
- Bearer-token authentication using Kubernetes Secrets.
- Integration with the new
configFromconfiguration mechanism. - Validation of conflicting or incompatible metrics settings.
- Isolated per-pod metrics storage to avoid multiprocess file collisions.
ServiceMonitor and metrics collection are disabled by default.
Related improvements: [#107](#107), [#110](#110), [#117](#117), [#118](#118).
User Experience Improvements
The web interface received several navigation and usability improvements.
Navigation
- Reorganized the sidebar into Operations, On-call, Configuration, and Administration sections.
- Added collapsible navigation groups.
- Sidebar group states are preserved across page reloads.
- Improved menu consistency, spacing, and localization.
Alerts and Incidents
- Improved AlertGroup Details layout and lifecycle action visibility.
- Made status, assignment, and escalation information easier to find.
- Added searchable AlertGroup selection when linking AlertGroups to Incidents, replacing manual ID entry.
- Improved search behavior, exact ID matching, and empty-result handling.
- Improved visibility of effective policies and their configuration sources.
Event Orchestration
- Improved execution trace presentation.
- Added more readable rule and action results.
- Added structured before/after comparisons of execution changes.
- Improved visualization of routing and policy selection.
Reliability and Bug Fixes
- Fixed Docker configuration changes not being applied correctly after restarts.
- Improved Prometheus counter aggregation across multiple Gunicorn workers.
- Fixed recent notification metrics to account for updates to existing delivery records.
- Improved database and migration health reporting during outages.
- Hardened metrics initialization when metrics are disabled or filesystem permissions are restricted.
- Improved metrics cleanup and preservation of historical process counters.
- Added validation for conflicting Helm configuration sources.
- Expanded regression tests for analytics, metrics, configuration, Docker, Helm, and localization.
Upgrade Notes
- Application version updated to 2.3.1.
- Helm chart version and
appVersionupdated to 2.3.1. - A database migration adds indexes used by recent notification metrics.
- Run the normal database migrations when upgrading.
- Prometheus metrics and Helm ServiceMonitor are opt-in.
- Existing configuration files remain supported.
- Preserve existing application encryption keys during upgrades.
- Restart web, scheduler, and worker processes after enabling or changing metrics settings.
Prometheus note: Multiprocess counters are aggregated within a shared process namespace, not automatically across separate Docker containers or Kubernetes pods. Database-backed gauges provide cross-component visibility.