Keeping Renewal Boring
Short-lived certificates are only a good idea if renewal is reliably automated. The price of short lifetimes is that expiry must be a non-event, and the only way to get there is automation.
Certificates with the setup presented live 90 days at most, usually less. Across a dozen hosts, each with a server certificate and one or more client certificates, manual renewal is tedious, and risks that something will expire unnoticed, and take a service down. The goal is a system where certificates are renewed before expiry, and the only time a human is involved is when renewal fails.
Per-host renewal
Each host runs a renewal service that watches its own certificates and renews them as they approach expiry. As a systemd unit:
[Unit]
Description=Auto-renew internal host certificate
After=network.target
[Service]
ExecStart=/usr/bin/step ca renew \
/etc/ssl/certs/service.crt \
/etc/ssl/private/service.key \
--daemon
Restart=always
RestartSec=10
User=root
StandardOutput=syslog
StandardError=syslog
SyslogIdentifier=step-renew
[Install]
WantedBy=multi-user.target
The important pieces:
--daemon: step runs continuously, waking on its own schedule to check whether the certificate is close enough to expiry to warrant renewal. It renews in place, so the certificate and key files are refreshed without intervention.Restart=always: if the renewal daemon dies, systemd brings it back. A renewal daemon that has quietly stopped is the silent failure this whole system exists to avoid.SyslogIdentifier=step-renew: renewal events are tagged in syslog. That tag is what makes renewal auditable so the logs can forwarded and can be reviewed with alerts triggered on failure.
The WAF's client certificates
The WAF holds a client certificate for every backend it talks to (waf-to-service, one per host). These are renewed on a weekly timer that reissues anything approaching expiry:
- A service performs the renewal pass, reissuing any
waf-to-<host>certificate that expires within the next several days. - A timer runs it weekly.
The specific window matters: renewing anything that expires "within the next 8 days" on a weekly schedule means every certificate gets multiple renewal attempts before it can possibly lapse. If one weekly run is missed for some reason, the next still catches the certificate with days to spare. The overlap is deliberate, you never want renewal timing to be so tight that a single missed run causes an outage.
Why the overlap and the logging matter
Two design choices on this page are really about one principle: renewal should fail loudly and early, not silently and late.
- Renew well before expiry, not at the last moment, so there's slack to notice and fix a problem before it becomes an outage. A certificate that renews only in its final hours has no margin for a failed run.
- Log the renewals and forward the logs, so a renewal that stops happening is detectable. The failure mode that actually hurts isn't a certificate expiring, it's a renewal daemon that silently stopped weeks ago, which you discover only when the certificate finally lapses. Monitoring the renewal, not just the expiry, is what turns that from a surprise outage into a log entry you can act on.
This is the part people skip, and it's the part that determines whether a private CA is a security improvement or a recurring self-inflicted outage. Short-lived certificates without reliable renewal are worse than long-lived ones. The automation is not optional infrastructure around the CA; it is the CA being usable.
Adapt this for…
Any certificate with a short lifetime, from any issuer. The principle is issuer-agnostic: automate renewal, give it comfortable margin before expiry, and monitor the renewal process itself rather than only the certificate's expiry date. Whether it's step-ca internal certs or public ACME certificates, the failure that hurts is always the same, renewal that stopped without anyone noticing, and the defense is always the same: make the renewal observable and give it slack.