Skip to main content

Monitor flag evaluation health with real-time metrics, error rate tracking, and auto-remediation

Last updated October 3, 2026

Flag Health

Flaggr tracks evaluation metrics for every flag, giving you real-time visibility into how flags are performing. The health system detects anomalies — high error rates, sudden drops in evaluation volume, or latency spikes — and surfaces them as health statuses.

Health Statuses

Every flag has a computed health status based on its recent evaluation metrics:

StatusCriteriaWhat it means
healthyError rate < 1%, normal evaluation volumeEverything is working as expected
warningError rate 1–5%, or evaluation volume anomalySomething may need attention
criticalError rate > 5%, or zero evaluations when expectedLikely a problem — investigate

Flag Health API

Get Health for a Single Flag

curl /api/flags/checkout-v2/health?serviceId=svc-web&projectId=proj-123

Response:

{
  "flagKey": "checkout-v2",
  "serviceId": "svc-web",
  "status": "healthy",
  "summary": {
    "totalEvaluations": 15420,
    "errorCount": 12,
    "errorRate": 0.0008,
    "avgLatencyMs": 2.3,
    "p95LatencyMs": 5.1,
    "p99LatencyMs": 8.7,
    "lastEvaluatedAt": "2026-02-19T14:32:01.000Z"
  },
  "reasonBreakdown": {
    "TARGETING_MATCH": 8200,
    "VARIANT": 4100,
    "DEFAULT": 3100,
    "ERROR": 12,
    "DISABLED": 8
  },
  "timeSeries": [
    {
      "timestamp": "2026-02-19T14:00:00.000Z",
      "evaluations": 450,
      "errors": 1,
      "avgLatencyMs": 2.1
    }
  ]
}

Get Health for All Flags in a Service

curl /api/services/svc-web/health?projectId=proj-123

Response:

{
  "serviceId": "svc-web",
  "overallStatus": "warning",
  "flags": [
    { "flagKey": "checkout-v2", "status": "healthy", "evaluations": 15420, "errorRate": 0.0008 },
    { "flagKey": "search-v3", "status": "warning", "evaluations": 890, "errorRate": 0.034 },
    { "flagKey": "onboarding-flow", "status": "critical", "evaluations": 0, "errorRate": 0 }
  ]
}

The overallStatus reflects the worst status among all flags — if any flag is critical, the service status is critical.

Evaluation Metrics Pipeline

Flaggr collects metrics for every flag evaluation:

MetricDescription
Evaluation countTotal evaluations per time bucket
Error countEvaluations that returned ERROR reason
Error rateErrors / total evaluations
LatencyEvaluation duration in milliseconds
Reason breakdownCount per evaluation reason (TARGETING_MATCH, VARIANT, DEFAULT, etc.)

Metrics are bucketed into 1-minute intervals and retained for 24 hours. Older data is automatically evicted.

Time Series Data

The health API returns time-series data for visualization. Each data point represents one minute:

{
  "timeSeries": [
    { "timestamp": "2026-02-19T14:00:00Z", "evaluations": 450, "errors": 1, "avgLatencyMs": 2.1 },
    { "timestamp": "2026-02-19T14:01:00Z", "evaluations": 462, "errors": 0, "avgLatencyMs": 1.9 },
    { "timestamp": "2026-02-19T14:02:00Z", "evaluations": 448, "errors": 3, "avgLatencyMs": 4.2 }
  ]
}

Use this data to plot evaluation volume, error rates, and latency over time in dashboards.

Percentile Calculations

The health API computes latency percentiles from the raw evaluation data:

PercentileDescription
p50Median latency — typical user experience
p9595th percentile — most users are faster than this
p9999th percentile — tail latency

Integrating Health Checks with Rollouts

Progressive rollout safety checks can reference the same metrics that the health API tracks:

{
  "safetyChecks": [
    { "type": "error_rate", "threshold": 0.05, "action": "rollback" },
    { "type": "latency_p99", "threshold": 500, "action": "pause" }
  ]
}

When you report health metrics to a rollout plan, the rollout engine compares them against your safety checks and automatically pauses or rolls back if thresholds are exceeded.

See Progressive Rollouts for details.

Monitoring Best Practices

  • Set up alerting for flags with critical or warning status
  • Watch for zero-evaluation flags — a flag that stops being evaluated may indicate a deployment issue
  • Monitor error rate trends, not just absolute values — a spike matters more than a steady 0.1%
  • Review stale flags — flags that haven't been evaluated in days may be candidates for cleanup
  • Use the service health endpoint for dashboards — it gives you a single view of all flags

Post-Toggle Safety & Drift Diagnostics

When a feature flag is toggled on or off, downstream systems may experience unexpected errors, latency regressions, or volume shifts. Flaggr includes an automated post-toggle drift diagnostic engine that continuously compares traffic before and after each state change.

How Drift Detection Works

  1. Audit Log Correlation: Flaggr indexes every toggle operation in Neon Postgres audit logs with microsecond precision, capturing the transition (from -> to), the acting engineer or service token, target environment, and propagation latency.
  2. Symmetric Time Window: The engine slices evaluation telemetry into symmetric windows (e.g. ±15m or ±30m) immediately preceding and following the toggle event.
  3. Delta Metric Computation:
    • Error Rate Shift (errorRateDelta): after.errorRate - before.errorRate. A surge indicates newly introduced runtime exceptions.
    • P99 Latency Shift (latencyP99DeltaMs): after.p99LatencyMs - before.p99LatencyMs. Identifies database query saturation or slow third-party API dependencies.
    • Volume Ratio (evaluationsRatio): after.evaluations / before.evaluations. Detects dropped requests or cascading retries.
  4. Automated Health Assessment:
    • healthy: No statistically significant error rate or latency regression.
    • warning: Error rate increased slightly (0.5%–2%) or P99 latency increased by >100ms.
    • degraded: Error rate jumped by >= 2% or went from 0% to > 1%. Immediate rollback is recommended.
    • insufficient_data: Fewer than 3 evaluations were recorded before/after the toggle.

Inspecting Post-Toggle Safety via API

curl -H "Authorization: Bearer fgr_token" \
  "https://flaggr.dev/api/flags/checkout-v2/toggle-impact?serviceId=web-app&environment=production"

Inspecting Post-Toggle Safety via MCP

AI agents running in Cursor, Claude Desktop, or CI pipelines can invoke:

{
  "tool": "get_flag_toggle_impact",
  "arguments": {
    "flagKey": "checkout-v2",
    "serviceId": "web-app",
    "environment": "production"
  }
}

Or trigger the guided verification workflow: verify_toggle_safety

Automated Guardrail Circuit Breaking & Auto-Rollback

When post-toggle metrics degrade beyond acceptable thresholds, Flaggr's circuit breaker triggers an automated rollback to protect downstream users:

POST /api/flags/checkout-v2/auto-rollback
Content-Type: application/json
 
{
  "serviceId": "web-app",
  "environment": "production",
  "dryRun": false,
  "maxErrorRateDelta": 0.02,
  "maxLatencyDeltaMs": 200
}
  • Reversion Mechanism: Automatically retrieves the immutable version snapshot preceding the toggle from flaggr.flag_versions and restores the exact state.
  • Audit Logging: Emits an audit log event with action flag.auto_rollback and delta metrics.
  • Cache Invalidation & SSE: Flushes Redis L2 caches and broadcasts real-time SSE stream events across connected SDK clients.
  • Multi-Channel Alerting: Dispatches urgent notifications to all configured organization and project alert channels (Slack, Discord, PagerDuty, Email, Webhook).

AI agents can verify or trigger auto-rollbacks via the MCP tool:

{
  "tool": "check_and_rollback_flag",
  "arguments": {
    "flagKey": "checkout-v2",
    "serviceId": "web-app",
    "environment": "production",
    "dryRun": true
  }
}

Flag Lifecycle & Automated Staleness Cleanup

Over time, feature flags accumulate and become technical debt. Flaggr's staleness engine continuously monitors evaluation volume and targeting stability to classify flags into lifecycle categories:

CategoryCriteriaActionable Recommendation
Stale Rolled-Out100% rolled out for $\ge 30$ daysSafe to remove: Remove conditional code branch and SDK calls from repository.
Stale DisabledDisabled for $\ge 30$ days with 0 evaluationsReady to archive: Archive or delete flag; remove dead fallback code.
Dead Flag0 evaluations across all environments for $\ge 60$ daysAbandoned flag: Safe to delete flag configuration immediately.
Launched100% rolled out for $\ge 7$ daysMonitoring stable: Plan repository cleanup in next sprint cycle.
Healthy ActiveActive evaluations with ongoing trafficKeep in codebase: Ongoing progressive rollout or experimentation.

Staleness API & AI Agent Integration

Query stale flags for an entire project:

GET /api/projects/{projectId}/stale-flags

Or invoke via MCP:

{
  "tool": "get_stale_flags",
  "arguments": {
    "projectId": "proj-production-web"
  }
}

The response provides computed staleness scores (0–100) and automated code removal advice for every flag eligible for retirement.