Observability overview
What each layer of an API7 Gateway deployment emits — metrics, logs, traces, and health endpoints — and which of them the product provides versus which you supply.
An API7 Gateway deployment has four layers that fail independently: the data plane that proxies traffic, the control plane that configures it, the PostgreSQL database that stores that configuration, and the Dashboard that people log in to. They do not emit the same signals, and API7 does not emit all of them — some come from your own platform. This page says which is which, so you know what you still have to instrument.
What each layer emits
| Layer | Metrics | Logs | Traces | Health endpoint |
|---|---|---|---|---|
| Data plane (gateway) | apisix_* families on 9091/apisix/prometheus/metrics, once the prometheus (opens in Plugin Hub docs) plugin runs as a global rule. See Monitor Metrics. | Access log and error log, format configurable. See Configure Centralized Logging. | Per-request spans via the opentelemetry (opens in Plugin Hub docs) plugin. See Configure Distributed Tracing. | 7085/status (liveness), 7085/status/ready (readiness — fails when the DP Manager is unreachable). Disabled by default — see Ports and Endpoints. |
| Control plane (Dashboard, DP Manager) | Go runtime metrics and go_sql_stats_connections_* (the PostgreSQL pool) on the status port at /metrics, plus product-specific series such as api7_dashboard_requests_total/api7_dashboard_request_duration_seconds and a full embedded-etcd family (etcd_server_*, etcd_mvcc_*, etcd_disk_*) from the DP Manager's etcd-compatible layer. See Metrics Catalogue. | Structured logs to stdout, level set by log.level in each component's conf.yaml. | Not emitted. | Dashboard 7081/healthz, DP Manager 7901/healthz — both bound to 127.0.0.1 by default. (The path is /healthz, not /status — a request to /status on either port returns 404.) |
| PostgreSQL | Not emitted by API7. Run postgres_exporter or use your managed provider's metrics. | Your PostgreSQL deployment's own logs. | Not emitted. | Your PostgreSQL deployment's own probes. |
| Dashboard / Admin API | Covered by the control-plane row — the Dashboard and the Admin API are the same process. | Access log to stdout, log.access_log in dashboard_conf/conf.yaml. | Not emitted. | 7081/healthz |
Two consequences worth reading twice:
- Database observability is yours to build. API7 stores every route, service, consumer and certificate in PostgreSQL, but it exports no database metrics. If you want to know about connection-pool exhaustion, replication lag or disk pressure before they take the control plane down, you have to instrument PostgreSQL yourself.
- Only the data plane is traced. A slow Dashboard save produces no span. Traces answer questions about proxied traffic, not about configuration operations.
Metrics
Data-plane metrics come from the prometheus plugin and are exposed on port 9091. Monitor Metrics covers enabling the plugin and the scrape configuration; the prometheus plugin reference (opens in Plugin Hub docs) has the full metric and label list.
Two operational cautions:
- The metrics depend on a
prometheusglobal rule. If your declarative configuration omitsglobal_rules.prometheus, anadc syncdeletes it and all gateway metrics stop. - Labels like
route,consumerandmatched_urimultiply series. Review cardinality against your Prometheus capacity before enabling them in a large deployment.
If you already run Prometheus, Use an Existing Prometheus covers both migration paths — scraping data planes directly, or receiving metrics through the DP Manager's remote-write channel.
Logging
The gateway writes an access log and an error log. Configure Centralized Logging documents the default format string, the JSON variant, and where each component writes.
Shipping them off the node:
- On Kubernetes, Collect Gateway Logs on Kubernetes covers an OpenTelemetry Collector DaemonSet, and Send Kubernetes Error Logs to Splunk covers error-log forwarding.
- Send Access Logs to Splunk uses the gateway's own logger plugin instead of a collector.
- Include Consumer Labels in Access Logs adds tenant or team attribution to each line, which is what makes per-consumer chargeback and per-team error budgets possible.
To correlate a log line with a trace, add $opentelemetry_trace_id and $opentelemetry_span_id to the log format — Configure Distributed Tracing shows the format string.
Tracing
The opentelemetry plugin emits one span per request with the route and status attached, and exports over OTLP. Configure Distributed Tracing covers the collector address, sampling, and the extra attributes worth attaching.
For a single misbehaving route, a trace backend is often more machinery than you need. Capture Request Traces with Debug Sessions records matching requests — including per-plugin and per-phase timings and the request's own log lines — from the Dashboard, without a collector.
Alerts and SLOs
Two mechanisms, and you will probably want both:
- The Dashboard's built-in alerting evaluates control-plane events — an instance going offline, a certificate approaching expiry, a licence quota being crossed. These are things Prometheus cannot see. Start at Configure Alerts.
- Prometheus alerting covers request-level symptoms: error ratio, latency, saturation. Alerts and SLOs has the SLIs, the PromQL, and how to choose your own thresholds.
Incident Dashboards has the panel queries for triage.
Related
- Ports and Endpoints — every listener named here, and what should reach it.
- Configure Readiness and Liveness Probes — wiring the health endpoints into Kubernetes or a load balancer.
- Troubleshoot API7 Gateway — turning these signals into a diagnosis.