Rotate Control-Plane and Data-Plane Certificates
Find out when a data plane's mTLS certificate expires, issue a replacement, and roll it out without dropping traffic.
Every data plane authenticates to the control plane with a client certificate issued by the control plane's CA. If one expires, that node stops receiving configuration. Issuance is automatic; rotation is not — nothing rotates these for you, and by default they are valid for ten years, so nothing will remind you either.
This guide covers finding the expiry dates, issuing replacements, and rolling them out with no downtime.
This is about the CP↔DP mTLS certificates, which secure configuration distribution. Certificates your clients see on 9443 are a separate concern — see SSL Certificates.
Why this is safe to do during traffic
Three properties, each verified against a running deployment, make rotation a rolling operation rather than a maintenance window:
- Issuing a certificate does not disturb anything. Nodes running on older certificates stay
Healthyand keep proxying while a new one is issued. - Certificates overlap. The control plane accepts every unexpired certificate its CA signed. Two data planes holding two different client certificates both connect and both serve traffic.
- Each node is independent. Replacing the certificate on one node has no effect on the others.
So you never need a coordinated cutover. Issue, then replace node by node.
Step 1: Find out when they expire
Each registered instance reports its own certificate expiry:
curl -k "https://localhost:7443/api/gateway_groups/${GATEWAY_GROUP}/instances" \
-H "X-API-KEY: ${API_KEY}"Every entry carries dataplane_certificate_expire_time as a Unix timestamp. To turn that into a list sorted by urgency:
curl -sk "https://localhost:7443/api/gateway_groups/${GATEWAY_GROUP}/instances" \
-H "X-API-KEY: ${API_KEY}" \
| jq -r '.list[]
| [ .hostname,
.status,
(.dataplane_certificate_expire_time | todate),
((.dataplane_certificate_expire_time - now) / 86400 | floor | tostring + "d")
] | @tsv' \
| sort -k3b41cede77cdb Healthy 2036-09-14T13:57:12Z 3650d
8294c4d21715 Healthy 2036-09-14T13:58:29Z 3650d3650d is the ten-year default. A node showing a much shorter window was installed with a shorter validity chosen at install time — see Mutual TLS between Control Plane and Data Plane.
Run this on a schedule. It is the only place the expiry is exposed per node, and the Dashboard's mTLS certificate between control plane and data plane will expire alert is the only thing that will tell you unprompted — configure it, at a daily interval, in Configure Alerts.
Step 2: Issue a replacement certificate
curl -k "https://localhost:7443/api/gateway_groups/${GATEWAY_GROUP}/dp_client_certificates" \
-X POST \
-H "X-API-KEY: ${API_KEY}" \
-H "Content-Type: application/json" \
-d '{}'The response carries the three values a data plane needs:
{
"value": {
"id": "688712c1-8cb1-4d88-950d-7e0bee1c8151",
"certificate": "-----BEGIN CERTIFICATE-----\n...",
"private_key": "-----BEGIN PRIVATE KEY-----\n...",
"ca_certificate": "-----BEGIN CERTIFICATE-----\n...",
"expiry": 2105012037,
"gateway_group_id": "default"
}
}Save them:
curl -sk "https://localhost:7443/api/gateway_groups/${GATEWAY_GROUP}/dp_client_certificates" \
-X POST -H "X-API-KEY: ${API_KEY}" -H "Content-Type: application/json" -d '{}' \
-o cert-response.json
jq -r '.value.certificate' cert-response.json > tls.crt
jq -r '.value.private_key' cert-response.json > tls.key
jq -r '.value.ca_certificate' cert-response.json > ca.crt
shred -u cert-response.json 2>/dev/null || rm -f cert-response.jsonThe response holds the private key, so remove the file once the three parts are extracted.
What changed: a new certificate exists. Nothing is using it yet, and every running node is unaffected.
The private key is returned once, in this response. It is not retrievable afterwards. If you lose it, issue another certificate.
Certificates are per gateway group. A node can only join the group whose CA signed its certificate.
Step 3: Roll it out, one node at a time
A data plane reads its certificate at startup, so the node has to restart to pick up a new one. Do one node at a time and let it come back Healthy before starting the next.
Update the Secret in place, then restart the workload:
kubectl create secret generic api7-ee-3-gateway-tls \
--from-file=tls.crt=./tls.crt \
--from-file=tls.key=./tls.key \
--from-file=ca.crt=./ca.crt \
-n api7 --dry-run=client -o yaml | kubectl apply -f -
kubectl -n api7 rollout restart deployment/api7-ee-3-gateway
kubectl -n api7 rollout status deployment/api7-ee-3-gateway --timeout=300sThe rollout replaces pods according to the Deployment's own strategy. With maxUnavailable: 0 and a readiness probe on /status/ready, replacement pods must be ready before old ones are removed, so capacity does not dip — see Configure Readiness and Liveness Probes.
Certificates arrive as environment variables, and a container's environment cannot be changed in place — recreate it:
docker rm -f api7-gateway-1
docker run -d --name api7-gateway-1 \
-e API7_DP_MANAGER_ENDPOINTS='["https://dp-manager:7943"]' \
-e API7_GATEWAY_GROUP_SHORT_ID="${GATEWAY_GROUP_SHORT_ID}" \
-e API7_DP_MANAGER_CERT="$(cat tls.crt)" \
-e API7_DP_MANAGER_KEY="$(cat tls.key)" \
-e API7_CONTROL_PLANE_CA="$(cat ca.crt)" \
-p 9080:9080 -p 9443:9443 \
api7/api7-ee-3-gateway:3.10.7Take the node out of your load balancer's pool before removing it, and put it back once it reports ready. A single-node deployment cannot rotate without an interruption; add a second node first.
Validate
After each node, confirm it came back on the new certificate:
curl -sk "https://localhost:7443/api/gateway_groups/${GATEWAY_GROUP}/instances" \
-H "X-API-KEY: ${API_KEY}" \
| jq -r '.list[] | [.hostname, .status, (.dataplane_certificate_expire_time | todate)] | @tsv'The rotated node reports Healthy with the new expiry date. Then confirm it is actually serving:
curl -i "http://<node>:9080/<a-known-route>"A node that registers as Healthy is a node whose mTLS handshake succeeded — that is what registration proves.
When every node shows the new expiry, the rotation is complete.
Roll back
Until a node has been restarted, it is still running on its previous certificate and there is nothing to roll back. If a restarted node fails to connect, put the previous tls.crt/tls.key/ca.crt back and restart it again — the old certificate is still valid, because nothing revoked it.
This is the reason to keep the previous certificate material until the whole fleet is rotated and verified, rather than deleting it as soon as the new one is issued.
What this does not cover
- Retiring an old certificate. There is no revocation or deletion API —
dp_client_certificatesonly issues. An old certificate stops being used when no node presents it any more, and stops being valid when it expires. If a key is compromised, that is a support matter, not a self-service operation. - Rotating the CA. The control plane generates its CA on first startup and this is not a self-service operation. Every client certificate chains to it, so it is not something to attempt from a runbook.
- Client-facing TLS. Certificates presented to API consumers on
9443are managed as SSL resources, not through this endpoint.
Related
- Mutual TLS between Control Plane and Data Plane — the trust model and the validity options
- Configure Alerts — the expiry alert that makes this proactive
- Run API7 Gateway in Production