Metrics Catalog
Every metric an API7 Gateway data plane exports, with its type, labels, and whether it depends on the prometheus global rule.
This is the set of metric families a data plane exposes on 9091/apisix/prometheus/metrics, captured from a running 3.10.7 gateway. Use it to size Prometheus, write queries against label names that exist, and know which series disappear when the prometheus global rule does.
For the plugin's own configuration and the meaning of each label value, see the prometheus plugin reference (opens in Plugin Hub docs). For enabling the endpoint, see Monitor Metrics.
The metrics listener binds to 127.0.0.1 by default. Publishing the port is not enough — set plugin_attr.prometheus.export_addr.ip to 0.0.0.0 first. The current 3.10.x release line also emits apisix_nginx_metric_errors_total twice from two registries, so Prometheus rejects a direct scrape. Use DP Manager remote write until your gateway includes a fix. See Monitor Metrics.
Traffic metrics
These come from the prometheus plugin running as a global rule. Delete that rule and these three families stop being updated, while everything in the next section keeps reporting normally. On a gateway that has already served traffic the series are not removed — they continue to be published at their last value until the data plane restarts, so the symptom is a flat line rather than a gap. See Manage Gateway Configuration with GitOps.
| Metric | Type | Labels |
|---|---|---|
apisix_http_status | counter | code, route, route_id, matched_uri, matched_host, service, service_id, consumer, node, gateway_group_id, instance_id, portal_id, api_product_id, request_type, request_llm_model, llm_model, mcp_request_type, mcp_tool_name, response_source |
apisix_http_latency | histogram | type, plus the same resource labels as above (no code, matched_uri, matched_host, response_source) |
apisix_bandwidth | counter | type, plus the same resource labels |
Notes that matter when writing queries:
apisix_http_latencyis in milliseconds, and exposes_bucket,_sumand_countseries. Quantiles needsum by (le)beforehistogram_quantile.typeonapisix_http_latencyisrequest,upstreamorapisix— client-observed, upstream wait, and the gateway's own overhead.typeonapisix_bandwidthisingressoregress.response_sourceis only onapisix_http_status, and is what separates a gateway fault (apisix,nginx) from an upstream fault (upstream).request_typeistraditional_http,websocket,ai_chatorai_stream. Excludewebsocketfrom latency queries — a WebSocket'srequestlatency is the lifetime of the connection.
Node metrics
Always exported, independent of the global rule.
| Metric | Type | Labels | What it tells you |
|---|---|---|---|
apisix_etcd_reachable | gauge | gateway_group_id, instance_id | 1 when this node can reach the DP Manager. 0 means it is serving stale configuration — it is not down, so availability stays green. |
apisix_etcd_modify_indexes | gauge | gateway_group_id, instance_id, key | Configuration revision per resource type. Divergence across instances indicates uneven rollout. |
apisix_node_info | gauge | gateway_group_id, instance_id, hostname | Always 1. Count it to get the number of instances reporting. |
apisix_shared_dict_capacity_bytes | gauge | gateway_group_id, instance_id, name | Configured size of each shared dictionary. |
apisix_shared_dict_free_space_bytes | gauge | gateway_group_id, instance_id, name | Remaining space. Exhaustion makes dict-backed plugins fail quietly — rate limiting stops counting rather than erroring. See Shared Memory Sizing. |
apisix_nginx_http_current_connections | gauge | gateway_group_id, instance_id, state | state is accepted, active, handled, reading, waiting or writing. |
apisix_http_requests_total | gauge | gateway_group_id, instance_id | NGINX's own request counter, independent of the plugin. |
apisix_nginx_metric_errors_total | counter | none | Errors inside the metrics subsystem itself. Non-zero means your metrics are unreliable. |
apisix_prometheus_disable | gauge | gateway_group_id, instance_id | Always present once the metrics endpoint is reachable, regardless of whether the prometheus global rule is enabled. Read the value, not the presence: 0 means the plugin is active and collecting traffic metrics, 1 means it is disabled. Do not alert on absent()/count() for this series — alert on the value instead. |
Upstream and stream metrics
These appear only once the corresponding feature is in use, so they are absent from a freshly installed gateway:
| Metric | Type | Labels | Appears when |
|---|---|---|---|
apisix_stream_connection_total | counter | listen_addr | Stream routes are in use. Connections handled per stream route. |
apisix_stream_status | counter | listen_addr, code, node | Stream routes are in use. Completed sessions by termination status. Added in 3.10.6. |
apisix_stream_active_connections · apisix_stream_bandwidth | gauge · counter | listen_addr, plus type and side on bandwidth | Stream routes are in use and the runtime is APISIX-Runtime. Added in 3.10.6. |
apisix_batch_process_entries | gauge | name, route_id, server_addr | A batching plugin (an HTTP or Kafka logger, for example) is configured. Remaining entries in its buffer. Registered lazily, on the first batch. |
apisix_ai_cache_hits_total · _misses_total · _bypasses_total · apisix_ai_cache_embedding_latency | counter · histogram | The full resource and model set, plus layer on _hits_total | The ai-cache plugin is configured. Added in 3.10.3. Labels are listed in full in AI Observability and Cost Tracking. |
apisix_stream_active_connections and apisix_stream_bandwidth are backed by a shared memory zone
that is only allocated on APISIX-Runtime. nginx_config.stream.metrics_zone_size defaults to
1m, which covers any realistic number of listening addresses, so the size is rarely what you
change — but on a runtime without the stream-metrics module the gateway logs that these metrics are
off and exports only apisix_stream_connection_total and apisix_stream_status.
Labels every data-plane series carries
gateway_group_id and instance_id are attached by API7 to every metric above. Prefer them over Prometheus's own instance target label when aggregating per node: instance_id is part of the exposition, so it survives both the direct-scrape path and the DP Manager remote-write path, while instance is added at scrape time and does not.
Cardinality
apisix_http_status, apisix_http_latency and apisix_bandwidth carry up to nineteen labels each. Series count multiplies across route × consumer × code × response_source × request_type, and response_source alone can triple an existing label combination.
Before enabling these across a large deployment, measure:
count({__name__=~"apisix_http_.*"})The plugin can drop labels it does not need — see the prometheus plugin's own configuration for which are optional. Dropping matched_uri and consumer is usually the largest single reduction.
Control-plane metrics
The Dashboard, DP Manager and Developer Portal each expose Prometheus metrics on their status port, at /metrics:
| Component | Status port | Default binding |
|---|---|---|
| Dashboard | 7081 | 127.0.0.1 |
| DP Manager | 7901 | 127.0.0.1 |
| Developer Portal | 4322 | 127.0.0.1 |
All three export the standard Go process metrics and the PostgreSQL pool metrics below, plus a set of product-specific series each component registers for itself:
| Metric | Type | Labels | What it tells you |
|---|---|---|---|
go_* | various | — | The standard Go runtime collector: go_threads, go_goroutines, go_memstats_*, go_gc_duration_seconds. Useful for a process that is leaking or stalling. |
go_sql_stats_connections_* | gauge · counter | db_name | The PostgreSQL connection pool. Each component registers one collector for its own pool, and db_name is always api7ee, so distinguish components by the scrape target rather than by this label. |
api7_dashboard_*, api7_dp_manager_*, api7_developer_portal_* | counter · histogram · summary | varies | Per-component RED metrics for the control plane's own HTTP layer — the Dashboard and DP Manager each additionally export requests_total and request_duration; all three export request/response size. This is the closest thing to a control-plane SLO signal this page has. |
etcd_* | various | varies | A full embedded-etcd metric family (etcd_server_*, etcd_mvcc_*, etcd_disk_*, and more — around 68 series) from the DP Manager's etcd-compatible layer. etcd_server_has_leader and etcd_server_leader_changes_seen_total are worth alerting on directly, since the DP Manager's etcd-compatible API depends on having a leader. |
os_fd_limit, os_fd_used | gauge | — | The process's file-descriptor ceiling and current usage — an early warning for FD exhaustion before it takes the component down. |
process_* | various | — | The standard process collector (CPU seconds, resident memory, open FDs), independent of os_fd_* above. |
Do not skip instrumenting the control plane because of the two rows above — there is a real, product-specific signal here, and etcd_server_has_leader/os_fd_used/the api7_* RED metrics are the ones to start with.
The pool series are the ones worth an alert, because they are the only direct view of the database contention that shows up as a slow Dashboard:
| Series | Meaning |
|---|---|
go_sql_stats_connections_max_open | The configured ceiling — database.max_open_conns in that component's conf.yaml, or the chart's max_open_conns. |
go_sql_stats_connections_open | Established connections, in use and idle. |
go_sql_stats_connections_in_use | Connections currently executing. |
go_sql_stats_connections_idle | Connections available for reuse. |
go_sql_stats_connections_waited_for | Cumulative count of requests that had to wait for a connection. |
go_sql_stats_connections_blocked_seconds | Cumulative time spent waiting. |
go_sql_stats_connections_closed_max_idle · _closed_max_lifetime · _closed_max_idle_time | Connections retired by each of the three pool limits. |
A rising waited_for with in_use pinned at max_open means the pool is the bottleneck, not PostgreSQL:
rate(go_sql_stats_connections_waited_for[5m]) > 0blocked_seconds climbing faster than wall-clock time means requests are queueing several deep.
Each component keeps its own pool against the same database, so raising max_open_conns on the Dashboard alone does not help if the DP Manager is the one waiting. Size them together against PostgreSQL's own max_connections, and remember the Developer Portal holds a third pool.
Because the status ports bind to 127.0.0.1 by default, an external Prometheus cannot reach them without changing server.status.host or scraping from inside the pod. These metrics are optional — the Dashboard does not consume them, and nothing in the product breaks if you never scrape them.
There is no gateway metric for upstream health-check state. The data plane reports health-check results to the control plane instead — read them from the Dashboard, or from GET /api/gateway_groups/{gateway_group_id}/services/{service_id}/healthcheck. $upstream_status exists as an NGINX variable and can be attached to apisix_http_status as an extra label, but it is the upstream's HTTP response code, not its health.
What is not here
- PostgreSQL's own metrics. API7 exports none. The
go_sql_stats_connections_*series above describe each client's pool, not the server: for table sizes, replication lag, checkpoint behaviour or lock waits, runpostgres_exporteror use your managed provider's metrics. - AI request metrics (
apisix_llm_*andapisix_ai_cache_*), which appear only for AI request types. See AI Observability and Cost Tracking for the full table.
Related
prometheusplugin reference (opens in Plugin Hub docs) — label semantics and plugin configuration- Monitor Metrics · Alerts and SLOs · Incident Dashboards