Alerts and SLOs
Service-level indicators for API7 Gateway with the PromQL to measure them, how to choose your own targets, and which alerts belong in Prometheus versus the Dashboard.
This page gives you the indicators, the queries, and a method for choosing targets. It deliberately does not hand you a set of thresholds to paste in: an availability target is a business decision about how much downtime your API consumers will tolerate, and a latency target depends on what your upstreams do. Numbers copied from a vendor page become alerts nobody trusts.
All queries below use metrics from the prometheus plugin (opens in Plugin Hub docs). Enable it first — see Monitor Metrics.
apisix_http_latency is measured in milliseconds, not seconds. A threshold written as 0.5 means half a millisecond.
Choosing targets
- Measure before you commit. Run each query below over 30 days of real traffic. Your current performance is the only honest starting point.
- Set the target just below what you already achieve. If you serve 99.95% successfully today, a 99.9% objective is defensible and leaves room to spend.
- Work out the error budget. 99.9% over 30 days is 43 minutes of failure. If that is obviously too much or absurdly little for your business, the target is wrong.
- Alert on burn rate, not on instantaneous breach. A single bad minute is noise. Burning a month's budget in an hour is an incident.
Service-level indicators
Availability — the share of requests the gateway did not fail
Separate failures the gateway caused from failures it faithfully relayed. The response_source label is what makes this possible: apisix means a gateway component produced the response (a plugin rejection, no route matched), nginx means the proxy could not complete the request (connection refused, upstream timeout), and upstream means your backend genuinely returned that status.
# Gateway-caused failure ratio — this is what your SLO should be written against.
sum(rate(apisix_http_status{code=~"5..", response_source=~"apisix|nginx"}[5m]))
/
sum(rate(apisix_http_status[5m]))# Upstream-caused failure ratio — your backends' problem, tracked separately.
sum(rate(apisix_http_status{code=~"5..", response_source="upstream"}[5m]))
/
sum(rate(apisix_http_status[5m]))Holding the gateway accountable for upstream 5xx makes the SLO unownable — the team that can fix it is not the team being paged.
Latency — how long the gateway itself added
apisix_http_latency carries a type label with three values: request is the full client-observed round trip, upstream is time spent waiting on your backend, and apisix is the difference — the gateway's own overhead.
# p99 of gateway overhead, in milliseconds. This is the gateway's SLI.
histogram_quantile(0.99,
sum by (le) (rate(apisix_http_latency_bucket{type="apisix", request_type!="websocket"}[5m]))
)# p99 of what the client experienced, for context.
histogram_quantile(0.99,
sum by (le) (rate(apisix_http_latency_bucket{type="request", request_type!="websocket"}[5m]))
)request_type!="websocket" matters from 3.10.7. A WebSocket session's request latency measures how long the upgraded connection stayed open, so a few long-lived sockets will drag any percentile to meaningless values.
Configuration freshness — can the data plane still reach the control plane
# 0 on any instance that has lost its DP Manager connection.
min by (instance_id) (apisix_etcd_reachable)Aggregate by (instance_id), not with a bare min() — a bare min() collapses the fleet to one number and throws away the label that tells you which node to look at. API7 attaches instance_id and gateway_group_id to every data-plane series, and those survive both scrape paths; Prometheus's own instance target label does not when metrics arrive through the DP Manager's remote-write channel.
An instance with a broken control-plane connection keeps serving traffic from its last known configuration. It is not down, so availability stays green — but it has silently stopped receiving changes. Alert on this separately, or a configuration rollout will appear to succeed while some fraction of your fleet ignores it.
Upstream health
# Upstream nodes currently marked unhealthy by active health checks.
count(apisix_upstream_status == 0) or vector(0)The or vector(0) matters on a dashboard: without it the filter returns an empty vector when every node is healthy, and the panel reads "No data" — indistinguishable from a broken scrape at exactly the moment you want reassurance.
Only reports on upstreams that have health checks configured — see Configure Upstream Health Checks.
Saturation
# Shared-dictionary utilization. Exhaustion breaks the plugins that depend on the dict
# — rate limiting stops counting, and it does so quietly.
1 - (apisix_shared_dict_free_space_bytes / apisix_shared_dict_capacity_bytes)# Active client connections per instance, against your configured worker limits.
sum by (instance_id) (apisix_nginx_http_current_connections{state="active"})Sizing guidance for the dictionaries is in Shared Memory Sizing.
Alerting on burn rate
Page when the error budget is being consumed fast enough to matter. The multiplier is how much faster than sustainable the budget is burning — at 14.4x, a 30-day budget is gone in about two days.
groups:
- name: api7-gateway-slo
rules:
# Replace 0.001 with (1 - your availability target).
- alert: API7GatewayErrorBudgetBurnFast
expr: |
(
sum(rate(apisix_http_status{code=~"5..", response_source=~"apisix|nginx"}[1h]))
/ sum(rate(apisix_http_status[1h]))
) > (14.4 * 0.001)
for: 2m
labels:
severity: page
annotations:
summary: "Gateway is burning its error budget 14.4x faster than sustainable"
- alert: API7GatewayControlPlaneUnreachable
expr: min by (instance_id) (apisix_etcd_reachable) == 0
for: 5m
labels:
severity: page
annotations:
summary: "Instance {{ $labels.instance_id }} has lost its DP Manager connection and is serving stale configuration"
- alert: API7GatewaySharedDictNearlyFull
expr: |
(1 - (apisix_shared_dict_free_space_bytes / apisix_shared_dict_capacity_bytes)) > 0.9
for: 15m
labels:
severity: ticket
annotations:
summary: "Shared dictionary {{ $labels.name }} is over 90% full"Validate the file before loading it:
promtool check rules api7-gateway-slo.rules.yamlWhat Prometheus cannot see
Some of the most consequential failures produce no request-level symptom until it is too late. The control plane knows about them and the Dashboard alerts on them directly — configure these in addition to the rules above, at Configure Alerts:
| Event | Why Prometheus misses it |
|---|---|
| mTLS certificate between control plane and data plane approaching expiry | Nothing is wrong yet. On expiry day every data plane loses its configuration channel at once. |
| Listener SSL certificate approaching expiry | Same shape, client-facing. |
| Gateway instance offline | An instance that stopped reporting also stopped being scraped. Absence of a series is not an error ratio. |
| Licence CPU quota exceeded | A licensing state, not a traffic signal. It can block configuration writes. |
| Healthy instance count below a floor | Requires knowing the expected count, which the control plane has and Prometheus does not. |
Related
- Monitor Metrics — enabling the plugin and the scrape target.
- Incident Dashboards — the panels to open when one of these fires.
prometheusplugin reference (opens in Plugin Hub docs) — every metric and label, including cardinality warnings.- Configure Alerts — the Dashboard's own alerting and contact points.