Docs
API7 GatewayReferenceReferenceMetrics Catalog

Metrics Catalog

Every metric an API7 Gateway data plane exports, with its type, labels, and whether it depends on the prometheus global rule.

This is the set of metric families a data plane exposes on 9091/apisix/prometheus/metrics, captured from a running 3.10.7 gateway. Use it to size Prometheus, write queries against label names that exist, and know which series disappear when the prometheus global rule does.

For the plugin's own configuration and the meaning of each label value, see the prometheus plugin reference (opens in Plugin Hub docs). For enabling the endpoint, see Monitor Metrics.

The metrics listener binds to 127.0.0.1 by default. Publishing the port is not enough — set plugin_attr.prometheus.export_addr.ip to 0.0.0.0 first. The current 3.10.x release line also emits apisix_nginx_metric_errors_total twice from two registries, so Prometheus rejects a direct scrape. Use DP Manager remote write until your gateway includes a fix. See Monitor Metrics.

Traffic metrics

These come from the prometheus plugin running as a global rule. Delete that rule and these three families stop being updated, while everything in the next section keeps reporting normally. On a gateway that has already served traffic the series are not removed — they continue to be published at their last value until the data plane restarts, so the symptom is a flat line rather than a gap. See Manage Gateway Configuration with GitOps.

MetricTypeLabels
apisix_http_statuscountercode, route, route_id, matched_uri, matched_host, service, service_id, consumer, node, gateway_group_id, instance_id, portal_id, api_product_id, request_type, request_llm_model, llm_model, mcp_request_type, mcp_tool_name, response_source
apisix_http_latencyhistogramtype, plus the same resource labels as above (no code, matched_uri, matched_host, response_source)
apisix_bandwidthcountertype, plus the same resource labels

Notes that matter when writing queries:

  • apisix_http_latency is in milliseconds, and exposes _bucket, _sum and _count series. Quantiles need sum by (le) before histogram_quantile.
  • type on apisix_http_latency is request, upstream or apisix — client-observed, upstream wait, and the gateway's own overhead.
  • type on apisix_bandwidth is ingress or egress.
  • response_source is only on apisix_http_status, and is what separates a gateway fault (apisix, nginx) from an upstream fault (upstream).
  • request_type is traditional_http, websocket, ai_chat or ai_stream. Exclude websocket from latency queries — a WebSocket's request latency is the lifetime of the connection.

Node metrics

Always exported, independent of the global rule.

MetricTypeLabelsWhat it tells you
apisix_etcd_reachablegaugegateway_group_id, instance_id1 when this node can reach the DP Manager. 0 means it is serving stale configuration — it is not down, so availability stays green.
apisix_etcd_modify_indexesgaugegateway_group_id, instance_id, keyConfiguration revision per resource type. Divergence across instances indicates uneven rollout.
apisix_node_infogaugegateway_group_id, instance_id, hostnameAlways 1. Count it to get the number of instances reporting.
apisix_shared_dict_capacity_bytesgaugegateway_group_id, instance_id, nameConfigured size of each shared dictionary.
apisix_shared_dict_free_space_bytesgaugegateway_group_id, instance_id, nameRemaining space. Exhaustion makes dict-backed plugins fail quietly — rate limiting stops counting rather than erroring. See Shared Memory Sizing.
apisix_nginx_http_current_connectionsgaugegateway_group_id, instance_id, statestate is accepted, active, handled, reading, waiting or writing.
apisix_http_requests_totalgaugegateway_group_id, instance_idNGINX's own request counter, independent of the plugin.
apisix_nginx_metric_errors_totalcounternoneErrors inside the metrics subsystem itself. Non-zero means your metrics are unreliable.
apisix_prometheus_disablegaugegateway_group_id, instance_idAlways present once the metrics endpoint is reachable, regardless of whether the prometheus global rule is enabled. Read the value, not the presence: 0 means the plugin is active and collecting traffic metrics, 1 means it is disabled. Do not alert on absent()/count() for this series — alert on the value instead.

Upstream and stream metrics

These appear only once the corresponding feature is in use, so they are absent from a freshly installed gateway:

MetricTypeLabelsAppears when
apisix_stream_connection_totalcounterlisten_addrStream routes are in use. Connections handled per stream route.
apisix_stream_statuscounterlisten_addr, code, nodeStream routes are in use. Completed sessions by termination status. Added in 3.10.6.
apisix_stream_active_connections · apisix_stream_bandwidthgauge · counterlisten_addr, plus type and side on bandwidthStream routes are in use and the runtime is APISIX-Runtime. Added in 3.10.6.
apisix_batch_process_entriesgaugename, route_id, server_addrA batching plugin (an HTTP or Kafka logger, for example) is configured. Remaining entries in its buffer. Registered lazily, on the first batch.
apisix_ai_cache_hits_total · _misses_total · _bypasses_total · apisix_ai_cache_embedding_latencycounter · histogramThe full resource and model set, plus layer on _hits_totalThe ai-cache plugin is configured. Added in 3.10.3. Labels are listed in full in AI Observability and Cost Tracking.

apisix_stream_active_connections and apisix_stream_bandwidth are backed by a shared memory zone that is only allocated on APISIX-Runtime. nginx_config.stream.metrics_zone_size defaults to 1m, which covers any realistic number of listening addresses, so the size is rarely what you change — but on a runtime without the stream-metrics module the gateway logs that these metrics are off and exports only apisix_stream_connection_total and apisix_stream_status.

Labels every data-plane series carries

gateway_group_id and instance_id are attached by API7 to every metric above. Prefer them over Prometheus's own instance target label when aggregating per node: instance_id is part of the exposition, so it survives both the direct-scrape path and the DP Manager remote-write path, while instance is added at scrape time and does not.

Cardinality

apisix_http_status, apisix_http_latency and apisix_bandwidth carry up to nineteen labels each. Series count multiplies across route × consumer × code × response_source × request_type, and response_source alone can triple an existing label combination.

Before enabling these across a large deployment, measure:

count({__name__=~"apisix_http_.*"})

The plugin can drop labels it does not need — see the prometheus plugin's own configuration for which are optional. Dropping matched_uri and consumer is usually the largest single reduction.

Control-plane metrics

The Dashboard, DP Manager and Developer Portal each expose Prometheus metrics on their status port, at /metrics:

ComponentStatus portDefault binding
Dashboard7081127.0.0.1
DP Manager7901127.0.0.1
Developer Portal4322127.0.0.1

All three export the standard Go process metrics and the PostgreSQL pool metrics below, plus a set of product-specific series each component registers for itself:

MetricTypeLabelsWhat it tells you
go_*various—The standard Go runtime collector: go_threads, go_goroutines, go_memstats_*, go_gc_duration_seconds. Useful for a process that is leaking or stalling.
go_sql_stats_connections_*gauge · counterdb_nameThe PostgreSQL connection pool. Each component registers one collector for its own pool, and db_name is always api7ee, so distinguish components by the scrape target rather than by this label.
api7_dashboard_*, api7_dp_manager_*, api7_developer_portal_*counter · histogram · summaryvariesPer-component RED metrics for the control plane's own HTTP layer — the Dashboard and DP Manager each additionally export requests_total and request_duration; all three export request/response size. This is the closest thing to a control-plane SLO signal this page has.
etcd_*variousvariesA full embedded-etcd metric family (etcd_server_*, etcd_mvcc_*, etcd_disk_*, and more — around 68 series) from the DP Manager's etcd-compatible layer. etcd_server_has_leader and etcd_server_leader_changes_seen_total are worth alerting on directly, since the DP Manager's etcd-compatible API depends on having a leader.
os_fd_limit, os_fd_usedgauge—The process's file-descriptor ceiling and current usage — an early warning for FD exhaustion before it takes the component down.
process_*various—The standard process collector (CPU seconds, resident memory, open FDs), independent of os_fd_* above.

Do not skip instrumenting the control plane because of the two rows above — there is a real, product-specific signal here, and etcd_server_has_leader/os_fd_used/the api7_* RED metrics are the ones to start with.

The pool series are the ones worth an alert, because they are the only direct view of the database contention that shows up as a slow Dashboard:

SeriesMeaning
go_sql_stats_connections_max_openThe configured ceiling — database.max_open_conns in that component's conf.yaml, or the chart's max_open_conns.
go_sql_stats_connections_openEstablished connections, in use and idle.
go_sql_stats_connections_in_useConnections currently executing.
go_sql_stats_connections_idleConnections available for reuse.
go_sql_stats_connections_waited_forCumulative count of requests that had to wait for a connection.
go_sql_stats_connections_blocked_secondsCumulative time spent waiting.
go_sql_stats_connections_closed_max_idle · _closed_max_lifetime · _closed_max_idle_timeConnections retired by each of the three pool limits.

A rising waited_for with in_use pinned at max_open means the pool is the bottleneck, not PostgreSQL:

rate(go_sql_stats_connections_waited_for[5m]) > 0

blocked_seconds climbing faster than wall-clock time means requests are queueing several deep.

Each component keeps its own pool against the same database, so raising max_open_conns on the Dashboard alone does not help if the DP Manager is the one waiting. Size them together against PostgreSQL's own max_connections, and remember the Developer Portal holds a third pool.

Because the status ports bind to 127.0.0.1 by default, an external Prometheus cannot reach them without changing server.status.host or scraping from inside the pod. These metrics are optional — the Dashboard does not consume them, and nothing in the product breaks if you never scrape them.

There is no gateway metric for upstream health-check state. The data plane reports health-check results to the control plane instead — read them from the Dashboard, or from GET /api/gateway_groups/{gateway_group_id}/services/{service_id}/healthcheck. $upstream_status exists as an NGINX variable and can be attached to apisix_http_status as an extra label, but it is the upstream's HTTP response code, not its health.

What is not here

  • PostgreSQL's own metrics. API7 exports none. The go_sql_stats_connections_* series above describe each client's pool, not the server: for table sizes, replication lag, checkpoint behaviour or lock waits, run postgres_exporter or use your managed provider's metrics.
  • AI request metrics (apisix_llm_* and apisix_ai_cache_*), which appear only for AI request types. See AI Observability and Cost Tracking for the full table.