Issuing and installing certificates
Pattern: bringing a new host into the trust domain. Every host follows the same bootstrap: make the internal names resolve, trust the CA root, then request the certificates it needs. Do these in order, the names have to resolve before certificates for those names can mean anything.
First: make the internal names resolve
Internal services are addressed by names like service.int.example. These names deliberately don't exist in public DNS, they're internal-only, so each host needs to resolve them locally. Rather than run an internal DNS server, this environment uses /etc/hosts entries, which is a right-sized choice for a fixed, small set of hosts.
On each host, /etc/hosts maps every internal service name to its private IP:
10.x.x.2 jump.int.example
10.x.x.4 pass.int.example
10.x.x.6 files.int.example
10.x.x.7 applications.int.example
10.x.x.9 auth.int.example
10.x.x.10 waf.int.example
# ...one line per host
This matters to mTLS directly, and it's a subtle failure mode worth understanding: mTLS verifies the certificate against the name being connected to. When the WAF connects to service.int.example, it checks the backend's certificate against that name. If /etc/hosts resolves the name to the wrong IP, or doesn't resolve it at all, you get failures that look like certificate problems, verification errors, connection refusals, but are actually name-resolution problems. When an mTLS connection misbehaves, confirm the name resolves to the right host before you start suspecting the certificate.
A gotcha while you're in here
Also map the host's own short hostname to the loopback range:
127.0.1.1 thishostname
Without it, sudo incurs a noticeable delay on every invocation, because it tries and fails to resolve the machine's own hostname before proceeding. The symptom (slow sudo) looks nothing like the cause (a missing hosts entry), which is exactly why it's worth writing down. It isn't strictly an mTLS concern, but every host gets this line as part of base configuration and it's the kind of thing you only debug once if you've seen it before.
Second: trust the CA root
Bootstrap the host's trust in the CA, pinning the root by the fingerprint captured when the CA was initialized:
sudo step ca bootstrap \
--ca-url https://10.x.x.2:9000 \
--fingerprint <CA_ROOT_FINGERPRINT>
Pinning by fingerprint is what makes this first trust decision safe. The host isn't blindly accepting whatever the CA presents; it's accepting only the specific root whose fingerprint you supplied out of band. Then install the root cert where the system and nginx will look for it:
sudo install -m 644 -D /root/.step/certs/root_ca.crt /etc/ssl/certs/example-ca.crt
From here on, this host trusts anything the CA signed and nothing it didn't.
Third: request the certificates
Recall the two-certificate model: each internal connection uses a server cert (held by the backend) and a client cert (held by the caller). For a backend host that serves an application, both get issued.
The host's server certificate, identifies the backend as itself:
sudo step ca certificate \
--san service.int.example \
--not-after 2160h --kty EC --curve P-256 \
service.int.example \
/etc/ssl/certs/service.crt \
/etc/ssl/private/service.key
The WAF's client certificate for reaching this host, issued on the WAF, identifies the WAF as the authorized caller:
sudo step ca certificate \
--not-after 2160h --kty EC --curve P-256 \
waf-to-service \
/etc/ssl/certs/waf-to-service.crt \
/etc/ssl/private/waf-to-service.key
A few things about these commands worth noting:
--not-after 2160hrequests the 90-day maximum. In practice the renewal automation reissues long before then; the explicit duration just pins it to the policy ceiling.--kty EC --curve P-256issues elliptic-curve certificates rather than RSA, smaller, faster, and entirely sufficient. A reasonable modern default.- The names are the identity.
service.int.exampleandwaf-to-servicearen't cosmetic; they're what the other end verifies against. This is why name resolution had to be sorted first.
A hard-won rule about shared certificates
One host in this environment serves several applications behind a single server certificate, one service.crt fronts multiple sites, distinguished by the request's host header rather than by separate certificates. That works well, but it carries a trap that is worth stating loudly because it causes a total, confusing outage:
Never reissue a shared server certificate just to add a name to it.
Reissuing that certificate to add a subject-alternative name for a new site doesn't append anything, it mints a brand-new certificate and key, and in doing so can disturb the chain the whole gateway depends on. The result is that every site on that host starts rejecting the mutually-authenticated connection at once, with the backend returning a TLS alert (access denied) and the proxy surfacing it as a misdirected-request error. It looks like the new site broke something; in fact the reissue broke the shared identity every site relies on.
The lesson, learned the hard way: on a host where one certificate serves many sites, the mTLS identity belongs to the host, not to any individual site. Sites are distinguished above the TLS layer, by host header. Adding a site means adding a route, never reissuing the host's certificate. If you find yourself about to reissue a shared cert to add a SAN, stop, there's almost certainly a routing-level way to add the site that doesn't touch the identity every other site depends on.
Adapt this for…
Any host joining the trust domain follows the same three steps: resolve the names, pin and install the CA root, request the certs it needs. A host that only makes outbound internal calls needs only a client cert; a host that serves internal traffic needs a server cert; a host that does both needs both. The pattern doesn't change with the number of hosts, which is what makes onboarding the twentieth host no harder than the second.