Docs

Survive and Recover a PostgreSQL Outage

What keeps working when the control plane's database fails, what silently does not, and the order of operations that brings it back.

API7 Gateway stores every route, service, consumer, certificate and policy in PostgreSQL. The product does not manage that database's availability for you — replication, promotion and backups are yours to run. What it does do is degrade in a specific and non-obvious way, and knowing the shape of that degradation is the difference between a database incident and an outage.

Everything below was verified against a running 3.10.7 deployment by stopping PostgreSQL.

What happens

ComponentDuring the outage
Data planes already runningKeep serving traffic normally. They proxy from configuration already in memory. This is the whole point of the split.
Dashboard — static pagesStill respond. Misleading: the UI loads.
Dashboard — login and any read or writeFail. Login returns 504 with The request exceeded the maximum allowed time limit.
Admin API and ADCFail for the same reason. No configuration change can be made.
A data plane that restartsCannot serve. It has no configuration to start from, and the DP Manager does not cache the last known configuration on its behalf. It logs failed to load the configuration: has no healthy etcd endpoint available and retries. This is expected — a data plane with no configuration has nothing to proxy.
DP ManagerAccepts TCP connections on 7943 but cannot answer while the database is down. Its logs fill with context canceled and 30-second timeouts on /v3/kv/range.

The dangerous part is not the outage. It is anything that restarts a data plane during it.

A node that restarts while the database is down does not serve again until the database is back, and what happens in between is not consistent. Two runs of the same procedure produced two different outcomes:

  • In one, the container stayed running and the gateway began serving on its own about four minutes after the database returned.
  • In the other, the process exited with status 1 — init_by_lua failing in the agent hook — and the container stayed exited. Nothing brought it back on its own.

Under Docker an exited container simply stays down until someone starts it. Under Kubernetes the kubelet restarts it, so the pod cycles until the database is back and then starts cleanly — visible as a climbing restart count or CrashLoopBackOff.

Either way the node is out of service for the rest of the outage plus a recovery delay, and you cannot predict which of the two you will get. So a rolling update, an autoscaler scaling in and back out, a spot-instance reclaim, or a node drain turns "traffic is unaffected" into lost capacity, at the moment you are least able to configure your way out of it.

During the outage

  1. Freeze anything that restarts data planes. Pause deployments, suspend the HorizontalPodAutoscaler, and hold off on node maintenance:

    kubectl -n api7 patch hpa gateway-hpa -p '{"spec":{"minReplicas":<current-replica-count>}}'
    kubectl -n api7 rollout pause deployment/api7-ee-3-gateway
  2. Do not restart surviving data planes to "fix" anything. They are the only thing still serving. Restarting one takes it out of service until the database is back.

  3. Accept that configuration is frozen. No route change, certificate upload or consumer edit is possible. Communicate this — teams will otherwise keep retrying.

  4. Recover the database, using your own PostgreSQL runbook. API7 has no part in this step.

After the database returns

Recovery is automatic. Give it a few minutes before intervening.

In a single-node test, the control plane accepted logins again within 30 seconds of PostgreSQL returning, the DP Manager's key-value layer stopped erroring in the same window, and a data plane that had failed to start during the outage began serving on its own about four minutes after the database came back. Nothing was restarted.

  1. Confirm PostgreSQL is genuinely ready, not merely running.

    kubectl -n api7 exec statefulset/api7-postgresql -- pg_isready -U api7ee
  2. Confirm the control plane recovered — log in, or call the Admin API. This should work almost immediately.

  3. Watch readiness and restart count, not container state. Depending on which outcome you hit, a node may be retrying inside a running container or may be exiting and being restarted.

    kubectl -n api7 get pods -l app.kubernetes.io/name=gateway \
      -o custom-columns='NAME:.metadata.name,READY:.status.containerStatuses[*].ready,RESTARTS:.status.containerStatuses[*].restartCount' -w
  4. Give it a few minutes before intervening. If the node is retrying internally, restarting it only starts the retry over. If it exited, Kubernetes is already restarting it for you; under Docker it needs docker start.

  5. Unfreeze what you froze during the outage.

Nothing has to be restarted for recovery to happen, including the DP Manager. In a single-node test with nothing touched, the control plane accepted logins again within 30 seconds of PostgreSQL returning and the data plane followed on its own. A restart during that window does not speed it up — it starts the retry over.

Validate

curl -sk "https://localhost:7443/api/gateway_groups/${GATEWAY_GROUP}/instances" \
  -H "X-API-KEY: ${API_KEY}" \
| jq -r '.list[] | [.hostname, .status] | @tsv'

Every instance should report Healthy. Then confirm a configuration write actually lands — a read succeeding is not proof the write path recovered:

adc dump -o post-incident.yaml

Finally, compare against the configuration in Git. If the database was restored from a backup rather than recovered in place, it may be behind what was deployed — see Manage Gateway Configuration with GitOps, which is what makes that comparison possible at all.

Reducing the blast radius beforehand

  • Configure a fallback control plane. Data Plane Resilience lets data planes bootstrap from object storage instead of the control plane, which removes the "a restart during an outage is fatal" failure mode. This is the single highest-value mitigation and it has to be in place before the incident.
  • Use /status/ready for readiness and /status for liveness. A node stuck at init_etcd passes liveness and fails readiness, so a correct probe configuration keeps it out of the load-balancer pool instead of blackholing traffic. See Configure Readiness and Liveness Probes.
  • Run PostgreSQL with replication and tested promotion. The gateway's tolerance buys you time; it does not remove the need.
  • Keep configuration in Git, so that recovering the database and recovering the configuration are separate problems.
  • Alert on apisix_etcd_reachable, which drops to 0 on a data plane that has lost the control plane — see Alerts and SLOs.