Docs

Incident dashboards

Triage panels for API7 Gateway, with the PromQL behind each one and the question it answers during an incident.

A triage dashboard exists to answer one question fast: is this the gateway, the upstream, or the control plane? The panels below are ordered so that you can answer it in about thirty seconds, then narrow down.

Every query uses metrics from the prometheus plugin (opens in Plugin Hub docs). Enable it first — see Monitor Metrics. For the alerts that bring you here, see Alerts and SLOs.

Row 1 — Is it us?

Put these four side by side at the top. Together they separate a gateway fault from an upstream fault, which is the decision that routes the incident to the right team.

PanelQueryRead it as
Request ratesum(rate(apisix_http_status[5m]))A cliff here means traffic stopped arriving — look upstream of the gateway, at DNS or the load balancer, not at the gateway.
Gateway-caused errorssum(rate(apisix_http_status{code=~"5..", response_source=~"apisix|nginx"}[5m]))Non-zero means the gateway or its proxying failed. Ours.
Upstream-caused errorssum(rate(apisix_http_status{code=~"5..", response_source="upstream"}[5m]))Your backend returned 5xx and the gateway relayed it faithfully. Theirs.
Gateway overhead p99histogram_quantile(0.99, sum by (le) (rate(apisix_http_latency_bucket{type="apisix", request_type!="websocket"}[5m])))Milliseconds the gateway added. Rising here with flat upstream latency points at a plugin.

The response_source label is what makes the second and third panels different. Without it, a backend outage and a gateway outage look identical on a 5xx graph, and the first twenty minutes of the incident get spent proving whose problem it is.

Row 2 — Where, exactly?

PanelQuery
Error rate by route (top 10)topk(10, sum by (route) (rate(apisix_http_status{code=~"5.."}[5m])))
Status code mixsum by (code) (rate(apisix_http_status[5m]))
Latency split — client vs upstream vs gatewayhistogram_quantile(0.99, sum by (le, type) (rate(apisix_http_latency_bucket{request_type!="websocket"}[5m])))
Unhealthy upstream nodesapisix_upstream_status == 0

The latency split panel is the one to keep. Three series — request, upstream, apisix — on one axis. request rising while upstream is flat means the gateway is the cause; both rising together means it is not.

topk by route is only useful if route labels are populated. route holds the route ID by default, or the route name when the plugin's prefer_name is true. Requests that matched no route carry an empty route label and group together — a large empty-label bucket in the 404 series usually means a routing change, not a client problem.

Row 3 — Is the fleet intact?

PanelQueryRead it as
Control-plane reachabilitymin by (instance_id) (apisix_etcd_reachable)Drops to 0 when any instance has lost the DP Manager. Those instances keep serving stale configuration — they look healthy while ignoring your changes.
Instances reportingcount(count by (instance_id) (apisix_node_info))Compare against your expected replica count. A drop here with flat request rate means the survivors absorbed the load.
Active connectionssum by (instance_id) (apisix_nginx_http_current_connections{state="active"})A single instance climbing while others are flat is a load-balancing problem, not a capacity problem.
Shared-dict utilization1 - (apisix_shared_dict_free_space_bytes / apisix_shared_dict_capacity_bytes)Above ~0.9, dict-backed plugins start failing quietly. Rate limiting stops counting rather than erroring.
Bandwidthsum by (type) (rate(apisix_bandwidth[5m]))ingress and egress in bytes/sec. A response-size change often explains latency that no latency panel accounts for.

What these panels cannot show you

Build the dashboard knowing its blind spots, so you do not conclude "everything is green" when it is not:

  • PostgreSQL. API7 exports no database metrics. Connection-pool exhaustion, replication lag and disk pressure are invisible here — add postgres_exporter panels or your managed provider's. A struggling database shows up on this dashboard only indirectly, as Dashboard slowness or failing configuration writes.
  • An instance that is completely gone. It stopped being scraped, so it contributes no error series. Use the instance-count panel, and the Dashboard's own "Gateway instance offline" alert.
  • Certificate expiry. No metric counts down to it. The control plane alerts on it — see Configure Alerts.
  • Configuration operations. Nothing here traces a Dashboard save or an adc sync.

Narrowing down to one request

Once the panels have told you which route is failing, Capture Request Traces with Debug Sessions records matching requests with per-plugin and per-phase timings and the request's own log lines — enough to identify which plugin added the latency the apisix series showed you.