Built-in alert rules for error rates, latency, gRPC health, and cache availability
Last updated October 3, 2026
Alerting
Flaggr includes built-in alert rules that monitor the health of the evaluation pipeline. Alerts are evaluated on every check and exposed via the /api/alerts endpoint.
Built-In Alert Rules
HighErrorRate
| Field | Value |
|---|---|
| Severity | Critical |
| Condition | Error rate exceeds 1% of total requests |
| Description | Evaluation errors are occurring at an abnormal rate |
Fires when the ratio of flaggr_evaluation_errors_total to flaggr_evaluations_total exceeds 0.01.
SlowP95Latency
| Field | Value |
|---|---|
| Severity | Warning |
| Condition | P95 request latency exceeds 100ms |
| Description | Flag evaluations are slower than expected |
Fires when the 95th percentile of flaggr_http_request_duration_seconds exceeds 100ms.
GrpcConnectionFailure
| Field | Value |
|---|---|
| Severity | Warning |
| Condition | Active gRPC errors detected |
| Description | gRPC streaming connections are failing |
Fires when gRPC connection errors are observed in the metrics.
CacheUnavailable
| Field | Value |
|---|---|
| Severity | Warning |
| Condition | Cache hit rate drops below 50% |
| Description | Cache layer is degraded or unavailable |
Fires when flaggr_cache_operations_total{result="hit"} / total cache operations falls below 0.5.
Querying Alerts
Get All Alerts
GET /api/alerts{
"status": "ok",
"firingCount": 0,
"alerts": [
{
"name": "HighErrorRate",
"status": "resolved",
"severity": "critical",
"description": "Error rate exceeds 1%",
"value": 0.002,
"threshold": 0.01
},
{
"name": "SlowP95Latency",
"status": "resolved",
"severity": "warning",
"description": "P95 latency exceeds 100ms",
"value": 45,
"threshold": 100
},
{
"name": "GrpcConnectionFailure",
"status": "resolved",
"severity": "warning",
"description": "gRPC connection errors detected",
"value": 0,
"threshold": 1
},
{
"name": "CacheUnavailable",
"status": "resolved",
"severity": "warning",
"description": "Cache hit rate below 50%",
"value": 0.95,
"threshold": 0.5
}
],
"evaluatedAt": "2025-07-20T10:00:00Z"
}Get Firing Alerts Only
GET /api/alerts?firing=trueReturns only alerts with status: "firing". The top-level status field reflects overall health:
"ok"— no alerts firing"alerting"— one or more alerts firing
Health Check Integration
The /api/health endpoint includes alert status. A load balancer or uptime monitor can check this endpoint and alert on "degraded" status.
GET /api/health{
"status": "healthy",
...
}Status values: "healthy", "degraded".
Prometheus Metrics
All alert rules are derived from the Prometheus metrics exposed at /api/metrics. You can build custom alerts in Grafana, Datadog, or any Prometheus-compatible monitoring tool:
# Error rate
sum(rate(flaggr_evaluation_errors_total[5m])) / sum(rate(flaggr_evaluations_total[5m]))
# P95 latency
histogram_quantile(0.95, sum(rate(flaggr_http_request_duration_seconds_bucket[5m])) by (le))
# Cache hit rate
sum(rate(flaggr_cache_operations_total{result="hit"}[5m])) / sum(rate(flaggr_cache_operations_total[5m]))
# Active gRPC connections
flaggr_grpc_active_connectionsNotification Channels & Providers
Flaggr supports multi-channel notification routing at both the Organization and Project levels:
| Provider | Mechanism | Features |
|---|---|---|
| Slack | Incoming Webhook | Rich Block Kit formatting, status color-coding, interactive links |
| Discord | Webhook | Embed cards with severity colors, metric fields, and timestamps |
| PagerDuty | Events API v2 | Automatic event deduplication, incident trigger, and custom metric details |
| Direct / Resend | Automated email dispatch with styled HTML status cards and recipients list | |
| Webhook | HTTP POST | HMAC-SHA256 signature verification (X-Flaggr-Signature), JSON payload |
Organization-Level Inheritance
Configure global alert channels once in Organization Settings (/console/org/[slug]/settings). All projects within the organization automatically inherit these channels without duplicating secrets or webhook URLs.
- Inherited channels are marked with an
Org-widebadge in project alert dashboards. - Project-level rules can target either project-specific channels or inherited organization channels.
Delivery Reliability Engine
Flaggr's AlertNotifier protects downstream endpoints with enterprise delivery semantics:
- Exponential Backoff Retries: Up to 3 attempts (1s, 2s, 4s delay).
- Per-Channel Circuit Breakers: Automatically opens after 3 consecutive failures to prevent cascade exhaustion; half-opens after 5 minutes.
- Time-Bucket Deduplication: Prevents notification storms when multiple serverless evaluation instances detect the same threshold breach.
- Dead-Letter Audit Trail: Records delivery failures in
flaggr.alert_delivery_failuresfor post-incident analysis.
Automated Guardrail Circuit Breaking & Auto-Rollback
When a flag toggle causes an unexpected regression in production, Flaggr's circuit breaker triggers immediate automated mitigation:
POST /api/flags/{flagKey}/auto-rollback
Content-Type: application/json
{
"serviceId": "svc-checkout",
"environment": "production",
"maxErrorRateDelta": 0.02,
"maxLatencyDeltaMs": 200
}- Drift Analysis: Compares symmetric pre- and post-toggle evaluation metrics.
- Automated Rollback: If health is
degraded(error surge $\ge 2%$ or P99 latency increase $\ge 200$ms), Flaggr reverts the flag to its previous version snapshot. - Emergency Alert Broadcast: Dispatches high-priority notifications across all configured organization and project channels.
Related
- Caching Strategy — Cache health and circuit breaker
- Rate Limiting — Rate limit monitoring
- Post-Toggle Safety & Drift Diagnostics — Post-toggle impact analysis
- REST API Reference — Alert endpoint details