Incident dashboards
Triage panels for API7 Gateway, with the PromQL behind each one and the question it answers during an incident.
A triage dashboard exists to answer one question fast: is this the gateway, the upstream, or the control plane? The panels below are ordered so that you can answer it in about thirty seconds, then narrow down.
Every query uses metrics from the prometheus plugin (opens in Plugin Hub docs). Enable it first — see Monitor Metrics. For the alerts that bring you here, see Alerts and SLOs.
Row 1 — Is it us?
Put these four side by side at the top. Together they separate a gateway fault from an upstream fault, which is the decision that routes the incident to the right team.
| Panel | Query | Read it as |
|---|---|---|
| Request rate | sum(rate(apisix_http_status[5m])) | A cliff here means traffic stopped arriving — look upstream of the gateway, at DNS or the load balancer, not at the gateway. |
| Gateway-caused errors | sum(rate(apisix_http_status{code=~"5..", response_source=~"apisix|nginx"}[5m])) | Non-zero means the gateway or its proxying failed. Ours. |
| Upstream-caused errors | sum(rate(apisix_http_status{code=~"5..", response_source="upstream"}[5m])) | Your backend returned 5xx and the gateway relayed it faithfully. Theirs. |
| Gateway overhead p99 | histogram_quantile(0.99, sum by (le) (rate(apisix_http_latency_bucket{type="apisix", request_type!="websocket"}[5m]))) | Milliseconds the gateway added. Rising here with flat upstream latency points at a plugin. |
The response_source label is what makes the second and third panels different. Without it, a backend outage and a gateway outage look identical on a 5xx graph, and the first twenty minutes of the incident get spent proving whose problem it is.
Row 2 — Where, exactly?
| Panel | Query |
|---|---|
| Error rate by route (top 10) | topk(10, sum by (route) (rate(apisix_http_status{code=~"5.."}[5m]))) |
| Status code mix | sum by (code) (rate(apisix_http_status[5m])) |
| Latency split — client vs upstream vs gateway | histogram_quantile(0.99, sum by (le, type) (rate(apisix_http_latency_bucket{request_type!="websocket"}[5m]))) |
| Unhealthy upstream nodes | apisix_upstream_status == 0 |
The latency split panel is the one to keep. Three series — request, upstream, apisix — on one axis. request rising while upstream is flat means the gateway is the cause; both rising together means it is not.
topk by route is only useful if route labels are populated. route holds the route ID by default, or the route name when the plugin's prefer_name is true. Requests that matched no route carry an empty route label and group together — a large empty-label bucket in the 404 series usually means a routing change, not a client problem.
Row 3 — Is the fleet intact?
| Panel | Query | Read it as |
|---|---|---|
| Control-plane reachability | min by (instance_id) (apisix_etcd_reachable) | Drops to 0 when any instance has lost the DP Manager. Those instances keep serving stale configuration — they look healthy while ignoring your changes. |
| Instances reporting | count(count by (instance_id) (apisix_node_info)) | Compare against your expected replica count. A drop here with flat request rate means the survivors absorbed the load. |
| Active connections | sum by (instance_id) (apisix_nginx_http_current_connections{state="active"}) | A single instance climbing while others are flat is a load-balancing problem, not a capacity problem. |
| Shared-dict utilization | 1 - (apisix_shared_dict_free_space_bytes / apisix_shared_dict_capacity_bytes) | Above ~0.9, dict-backed plugins start failing quietly. Rate limiting stops counting rather than erroring. |
| Bandwidth | sum by (type) (rate(apisix_bandwidth[5m])) | ingress and egress in bytes/sec. A response-size change often explains latency that no latency panel accounts for. |
What these panels cannot show you
Build the dashboard knowing its blind spots, so you do not conclude "everything is green" when it is not:
- PostgreSQL. API7 exports no database metrics. Connection-pool exhaustion, replication lag and disk pressure are invisible here — add
postgres_exporterpanels or your managed provider's. A struggling database shows up on this dashboard only indirectly, as Dashboard slowness or failing configuration writes. - An instance that is completely gone. It stopped being scraped, so it contributes no error series. Use the instance-count panel, and the Dashboard's own "Gateway instance offline" alert.
- Certificate expiry. No metric counts down to it. The control plane alerts on it — see Configure Alerts.
- Configuration operations. Nothing here traces a Dashboard save or an
adc sync.
Narrowing down to one request
Once the panels have told you which route is failing, Capture Request Traces with Debug Sessions records matching requests with per-plugin and per-phase timings and the request's own log lines — enough to identify which plugin added the latency the apisix series showed you.
Related
- Alerts and SLOs — the indicators and alert rules these panels support.
- Use an Existing Prometheus — getting these metrics into the Prometheus your dashboards already read.
- Troubleshoot API7 Gateway — from symptom to cause.