Monitor flag evaluation health with real-time metrics, error rate tracking, and auto-remediation
Last updated October 3, 2026
Flag Health
Flaggr tracks evaluation metrics for every flag, giving you real-time visibility into how flags are performing. The health system detects anomalies — high error rates, sudden drops in evaluation volume, or latency spikes — and surfaces them as health statuses.
Health Statuses
Every flag has a computed health status based on its recent evaluation metrics:
| Status | Criteria | What it means |
|---|---|---|
| healthy | Error rate < 1%, normal evaluation volume | Everything is working as expected |
| warning | Error rate 1–5%, or evaluation volume anomaly | Something may need attention |
| critical | Error rate > 5%, or zero evaluations when expected | Likely a problem — investigate |
Flag Health API
Get Health for a Single Flag
curl /api/flags/checkout-v2/health?serviceId=svc-web&projectId=proj-123Response:
{
"flagKey": "checkout-v2",
"serviceId": "svc-web",
"status": "healthy",
"summary": {
"totalEvaluations": 15420,
"errorCount": 12,
"errorRate": 0.0008,
"avgLatencyMs": 2.3,
"p95LatencyMs": 5.1,
"p99LatencyMs": 8.7,
"lastEvaluatedAt": "2026-02-19T14:32:01.000Z"
},
"reasonBreakdown": {
"TARGETING_MATCH": 8200,
"VARIANT": 4100,
"DEFAULT": 3100,
"ERROR": 12,
"DISABLED": 8
},
"timeSeries": [
{
"timestamp": "2026-02-19T14:00:00.000Z",
"evaluations": 450,
"errors": 1,
"avgLatencyMs": 2.1
}
]
}Get Health for All Flags in a Service
curl /api/services/svc-web/health?projectId=proj-123Response:
{
"serviceId": "svc-web",
"overallStatus": "warning",
"flags": [
{ "flagKey": "checkout-v2", "status": "healthy", "evaluations": 15420, "errorRate": 0.0008 },
{ "flagKey": "search-v3", "status": "warning", "evaluations": 890, "errorRate": 0.034 },
{ "flagKey": "onboarding-flow", "status": "critical", "evaluations": 0, "errorRate": 0 }
]
}The overallStatus reflects the worst status among all flags — if any flag is critical, the service status is critical.
Evaluation Metrics Pipeline
Flaggr collects metrics for every flag evaluation:
| Metric | Description |
|---|---|
| Evaluation count | Total evaluations per time bucket |
| Error count | Evaluations that returned ERROR reason |
| Error rate | Errors / total evaluations |
| Latency | Evaluation duration in milliseconds |
| Reason breakdown | Count per evaluation reason (TARGETING_MATCH, VARIANT, DEFAULT, etc.) |
Metrics are bucketed into 1-minute intervals and retained for 24 hours. Older data is automatically evicted.
Time Series Data
The health API returns time-series data for visualization. Each data point represents one minute:
{
"timeSeries": [
{ "timestamp": "2026-02-19T14:00:00Z", "evaluations": 450, "errors": 1, "avgLatencyMs": 2.1 },
{ "timestamp": "2026-02-19T14:01:00Z", "evaluations": 462, "errors": 0, "avgLatencyMs": 1.9 },
{ "timestamp": "2026-02-19T14:02:00Z", "evaluations": 448, "errors": 3, "avgLatencyMs": 4.2 }
]
}Use this data to plot evaluation volume, error rates, and latency over time in dashboards.
Percentile Calculations
The health API computes latency percentiles from the raw evaluation data:
| Percentile | Description |
|---|---|
| p50 | Median latency — typical user experience |
| p95 | 95th percentile — most users are faster than this |
| p99 | 99th percentile — tail latency |
Integrating Health Checks with Rollouts
Progressive rollout safety checks can reference the same metrics that the health API tracks:
{
"safetyChecks": [
{ "type": "error_rate", "threshold": 0.05, "action": "rollback" },
{ "type": "latency_p99", "threshold": 500, "action": "pause" }
]
}When you report health metrics to a rollout plan, the rollout engine compares them against your safety checks and automatically pauses or rolls back if thresholds are exceeded.
See Progressive Rollouts for details.
Monitoring Best Practices
- Set up alerting for flags with
criticalorwarningstatus - Watch for zero-evaluation flags — a flag that stops being evaluated may indicate a deployment issue
- Monitor error rate trends, not just absolute values — a spike matters more than a steady 0.1%
- Review stale flags — flags that haven't been evaluated in days may be candidates for cleanup
- Use the service health endpoint for dashboards — it gives you a single view of all flags
Post-Toggle Safety & Drift Diagnostics
When a feature flag is toggled on or off, downstream systems may experience unexpected errors, latency regressions, or volume shifts. Flaggr includes an automated post-toggle drift diagnostic engine that continuously compares traffic before and after each state change.
How Drift Detection Works
- Audit Log Correlation: Flaggr indexes every toggle operation in Neon Postgres audit logs with microsecond precision, capturing the transition (
from->to), the acting engineer or service token, target environment, and propagation latency. - Symmetric Time Window: The engine slices evaluation telemetry into symmetric windows (e.g. ±15m or ±30m) immediately preceding and following the toggle event.
- Delta Metric Computation:
- Error Rate Shift (
errorRateDelta):after.errorRate - before.errorRate. A surge indicates newly introduced runtime exceptions. - P99 Latency Shift (
latencyP99DeltaMs):after.p99LatencyMs - before.p99LatencyMs. Identifies database query saturation or slow third-party API dependencies. - Volume Ratio (
evaluationsRatio):after.evaluations / before.evaluations. Detects dropped requests or cascading retries.
- Error Rate Shift (
- Automated Health Assessment:
healthy: No statistically significant error rate or latency regression.warning: Error rate increased slightly (0.5%–2%) or P99 latency increased by >100ms.degraded: Error rate jumped by >= 2% or went from 0% to > 1%. Immediate rollback is recommended.insufficient_data: Fewer than 3 evaluations were recorded before/after the toggle.
Inspecting Post-Toggle Safety via API
curl -H "Authorization: Bearer fgr_token" \
"https://flaggr.dev/api/flags/checkout-v2/toggle-impact?serviceId=web-app&environment=production"Inspecting Post-Toggle Safety via MCP
AI agents running in Cursor, Claude Desktop, or CI pipelines can invoke:
{
"tool": "get_flag_toggle_impact",
"arguments": {
"flagKey": "checkout-v2",
"serviceId": "web-app",
"environment": "production"
}
}Or trigger the guided verification workflow:
verify_toggle_safety
Automated Guardrail Circuit Breaking & Auto-Rollback
When post-toggle metrics degrade beyond acceptable thresholds, Flaggr's circuit breaker triggers an automated rollback to protect downstream users:
POST /api/flags/checkout-v2/auto-rollback
Content-Type: application/json
{
"serviceId": "web-app",
"environment": "production",
"dryRun": false,
"maxErrorRateDelta": 0.02,
"maxLatencyDeltaMs": 200
}- Reversion Mechanism: Automatically retrieves the immutable version snapshot preceding the toggle from
flaggr.flag_versionsand restores the exact state. - Audit Logging: Emits an audit log event with action
flag.auto_rollbackand delta metrics. - Cache Invalidation & SSE: Flushes Redis L2 caches and broadcasts real-time SSE stream events across connected SDK clients.
- Multi-Channel Alerting: Dispatches urgent notifications to all configured organization and project alert channels (Slack, Discord, PagerDuty, Email, Webhook).
AI agents can verify or trigger auto-rollbacks via the MCP tool:
{
"tool": "check_and_rollback_flag",
"arguments": {
"flagKey": "checkout-v2",
"serviceId": "web-app",
"environment": "production",
"dryRun": true
}
}Flag Lifecycle & Automated Staleness Cleanup
Over time, feature flags accumulate and become technical debt. Flaggr's staleness engine continuously monitors evaluation volume and targeting stability to classify flags into lifecycle categories:
| Category | Criteria | Actionable Recommendation |
|---|---|---|
| Stale Rolled-Out | 100% rolled out for $\ge 30$ days | Safe to remove: Remove conditional code branch and SDK calls from repository. |
| Stale Disabled | Disabled for $\ge 30$ days with 0 evaluations | Ready to archive: Archive or delete flag; remove dead fallback code. |
| Dead Flag | 0 evaluations across all environments for $\ge 60$ days | Abandoned flag: Safe to delete flag configuration immediately. |
| Launched | 100% rolled out for $\ge 7$ days | Monitoring stable: Plan repository cleanup in next sprint cycle. |
| Healthy Active | Active evaluations with ongoing traffic | Keep in codebase: Ongoing progressive rollout or experimentation. |
Staleness API & AI Agent Integration
Query stale flags for an entire project:
GET /api/projects/{projectId}/stale-flagsOr invoke via MCP:
{
"tool": "get_stale_flags",
"arguments": {
"projectId": "proj-production-web"
}
}The response provides computed staleness scores (0–100) and automated code removal advice for every flag eligible for retirement.
Related
- REST API Reference — Endpoints for toggle-impact and metrics
- MCP Server Guide — Using Flaggr with Claude, Cursor, and AI agents
- Progressive Rollouts — Safety checks reference health metrics
- Alerting — Built-in health monitoring and notifications
- Performance — Latency benchmarks and optimization