Secret Rotation Is a Distributed-System Migration
How to introduce new credentials, coordinate consumers, and retire old dependencies safely
Today I looked into secret rotation. The part that needed the most thought was what happens after a replacement secret exists.
A running worker might still have the old value in memory. An API token issued months ago might still be valid. An encrypted database row might need the old key long after every application server has picked up the new one.
Those are dependencies to migrate. Updating a secret-store entry only handles one part of the work.
The useful general pattern is expand → switch → drain → retire. What changes between systems is how they drain their dependencies on the old secret. I’ll work through HMAC token fingerprints first, then compare signing keys, encryption keys, and credentials used to call another service.
Expand, switch, drain, retire
Before rotating anything, identify who uses the secret and what outlives the process using it. That might include API replicas, scheduled jobs, external clients, issued tokens, queued work, stored data, and backups.
Where the system supports overlap, the rollout looks like this:
Select any diagram to view it at full size.
The ordering matters. A consumer that may encounter newly produced work must be ready for the new version before a producer starts depending on it. For an outbound service credential, the receiving service must accept the replacement before callers start sending it.
Draining isn’t necessarily waiting a few minutes. A signed token may expire naturally. An encrypted record may need a migration job. A dormant API token may need its owner to replace it. Decide what counts as drained before beginning the rollout.
Some providers accept only one credential at a time. In that case, this pattern still identifies the coordination work, but a period where both versions work may require a second service identity or another provider-supported mechanism. Without that support, plan for retries or an interruption.
Case study: rotating an API token’s HMAC key
Which secret is changing?
There are two secrets in this example:
| Value | Where it lives | What it does |
|---|---|---|
| API token | Client; briefly in server memory during a request | Proves possession of the credential |
| HMAC key | Secret store and application memory | Computes the token’s stored fingerprint |
The database stores a fingerprint of the token, rather than the token itself:
Client holds: token
Database stores: HMAC-SHA256(key, token)
On a request, the server computes the fingerprint again and looks for a matching credential. It then checks expiration, revocation, the associated account, and permissions.
HMAC is a keyed hash, not encryption. There is no decrypt operation that gives us the token back. The construction is described in RFC 2104.
We’re rotating the HMAC key while keeping the client’s token unchanged. A fast HMAC isn’t a substitute for password hashing; the tokens here are generated with enough secure randomness to resist guessing.
Read the diagram from top to bottom. The blue token stays the same in all three requests. The database value changes from the old key’s fingerprint to the new key’s fingerprint.
Why replacing the key breaks existing tokens
Suppose a credential was stored using K1. After replacing the key, the same client token produces a different fingerprint:
Stored: HMAC(K1, token)
Looked up: HMAC(K2, token)
The client hasn’t changed anything, but the database lookup fails.
With multiple servers, this can turn into intermittent authentication failures. A server using K2 issues a token. The next request reaches a server that only knows K1, which can’t find the credential.
That gives us a rollout rule: every running process that validates tokens must accept every key that a running process can use to write fingerprints. That includes keys used when migrating a credential, not just when issuing one.
Give the application a key ring
Keep the accepted keys and the current write key as separate configuration:
Accepted keys: k1 -> K1, k2 -> K2
Write key ID: k1
k1 and k2 are identifiers. K1 and K2 are the secret bytes. Keep each ID permanently associated with the same bytes. Replacing the bytes behind k1 would make every existing k1 record misleading.
Store the ID with the fingerprint:
k1$<64-character SHA-256 hexadecimal digest>
You can also store the ID and digest in separate columns. Either way, make the full fingerprint searchable with a unique index, and leave room for the identifier if your existing column only fits the digest.
Here’s the small amount of Python needed to produce and look up those values:
import hmac
import secrets
def fingerprint(token: str, key_id: str, key: bytes) -> str:
digest = hmac.new(key, token.encode("utf-8"), "sha256").hexdigest()
return f"{key_id}${digest}"
def lookup_candidates(token: str, keys: dict[str, bytes]) -> list[str]:
return [
fingerprint(token, key_id, key)
for key_id, key in keys.items()
]
# Fresh keys for this local demonstration only.
keys = {"k1": secrets.token_bytes(32), "k2": secrets.token_bytes(32)}
write_key_id = "k1"
token = secrets.token_urlsafe(32)
stored = fingerprint(token, write_key_id, keys[write_key_id])
# Switch the writer; the old record remains discoverable.
write_key_id = "k2"
assert stored in lookup_candidates(token, keys)
migrated = fingerprint(token, write_key_id, keys[write_key_id])
assert migrated != stored
Python’s hmac and secrets modules provide these operations. In an application, provision each key once and load those same bytes from your secret store. Generating keys on each startup, as the local example does, would break authentication on every restart.
Validate the configuration before a process becomes ready: the write key must be present, IDs must be unambiguous, and key material must decode correctly. Keep the token encoding unchanged across versions. Take one consistent snapshot of the key ring for each request if configuration can reload while the process is running.
Find the credential before knowing its key ID
The client sends an opaque token. It doesn’t send the key ID stored in the database, so we can’t read that ID until we’ve found the row.
Compute a candidate fingerprint for each accepted key, then look for any of them in one parameterized query:
SELECT id, fingerprint, expires_at, revoked_at, principal_id
FROM credentials
WHERE fingerprint IN (:candidate_k1, :candidate_k2);
The placeholders are illustrative; bind them using your database driver. A match still needs the usual credential and authorization checks. A missing match means authentication fails.
If your existing records contain bare digests, include a bare digest under the legacy key as another candidate until those records are migrated. Deploy that compatibility support before writing the new prefixed format.
With two accepted keys, this approach computes two HMACs and makes one indexed lookup. Keep the key ring small. A growing list of forgotten keys adds work to every request and makes retirement harder.
Some token formats include a public credential ID, allowing the server to fetch the row first and verify its stored digest with one key. That is a different lookup design; if you compare digests in application code, use hmac.compare_digest.
Apply the four stages to HMAC fingerprints
Throughout expansion, switching, and draining, validators accept both K1 and K2. The write key changes from K1 to K2 only after that compatibility support is deployed.
1. Expand: distribute K2 while continuing to write with K1
Create K2 and deploy the expanded key ring. Keep K1 as the write key.
Verify the configuration loaded by the actual running processes. An updated secret-store entry or a completed CI job doesn’t prove that an old worker has refreshed its cached keys. Include API replicas, workers, scheduled jobs, and any tools that issue or validate credentials.
In this design, the application receives the key bytes and computes HMACs locally. A KMS that performs MAC operations without exporting the key is a separate architecture; swapping its key alias still doesn’t migrate your database fingerprints.
2. Switch: write new fingerprints with K2
Once every reader supports K2, change the write key and refresh the fleet.
During this rollout, some processes may still write with K1. That’s safe because every reader accepts both keys. Check both directions: a token issued by an old process must work on a new process, and a token issued by a new process must work on an old one.
Wait until all writers use K2 before enabling migration. Otherwise, a process whose current key is still K1 might rewrite a K2 record back to K1. Issuance and migration can use separate controls for this reason.
3. Drain: migrate a fingerprint when its token is used
When an old candidate matches a valid credential, the request supplies the token we need:
Before request: k1$HMAC(K1, token)
After request: k2$HMAC(K2, token)
The client keeps its token. Only the stored fingerprint changes. This is migration on use, sometimes called lazy migration. It works here because the request supplies an input that the database doesn’t retain.
Use a conditional update so migration doesn’t overwrite a concurrent replacement of the credential:
UPDATE credentials
SET fingerprint = :new_fingerprint
WHERE id = :credential_id
AND fingerprint = :observed_fingerprint
AND revoked_at IS NULL
AND (expires_at IS NULL OR expires_at > CURRENT_TIMESTAMP);
Bind the values, check the affected row count, and commit the update. Here, observed_fingerprint is the exact value returned by the lookup, and new_fingerprint is computed from the submitted token using K2.
Two requests may read the same old value. One updates it; the other affects zero rows. Re-read the credential and confirm that the submitted token still matches an accepted fingerprint and that the credential is still usable. Zero updated rows could also mean revocation, expiry, or replacement; it doesn’t automatically mean another request completed the migration.
The condition protects the migration write. Your application’s transaction and authorization rules still determine how concurrent revocation affects an in-flight request.
4. Retire: remove K1 after its remaining credentials are handled
Before removing K1, count the usable credentials still stored under it. Break that count down by credential type or owner so someone can act on the remaining dependencies. A query for prefixed records might look like:
SELECT COUNT(*)
FROM credentials
WHERE fingerprint LIKE 'k1$%'
AND revoked_at IS NULL
AND (expires_at IS NULL OR expires_at > CURRENT_TIMESTAMP);
Include legacy bare digests in your accounting if they still exist. With separate columns, filter on the key ID instead.
Once the remaining credentials have expired, migrated, been replaced, or been deliberately revoked, remove K1 from the accepted set and refresh every process holding it. Keep it out of configurations that autoscaling or rollback can launch later.
Dormant tokens need a decision
A batch job can’t convert HMAC(K1, token) into HMAC(K2, token) using only the old fingerprint and the two keys. It needs the original token.
For credentials that are used regularly, migration happens through normal traffic. For a token sitting in an integration that runs once a quarter, nothing happens until that integration sends a request.
If tokens have a fixed maximum lifetime, you can retire K1 after all K1-dependent credentials have expired or migrated. Count that lifetime from the last possible K1 write, including migrations, and account for any expiry tolerance your validator allows. If the same credential can be extended indefinitely, that timing argument no longer holds.
For tokens that never expire, choose how to handle the remaining records: keep K1 available, arrange replacement with their owners, or revoke them. Quiet traffic isn’t evidence that those tokens are safe to invalidate.
Old-key lookup traffic won’t tell you how many dormant credentials remain. Count the records too. Rehashing on token use is specific to this storage design; it isn’t a general step in every secret rotation.
The same rollout, different dependencies
Signing keys: let old tokens expire
For asymmetric signed tokens, verifiers need the new public key before the issuer starts signing with the new private key. Publish the new public key through your trusted key-distribution mechanism, such as a configured JWKS endpoint, and account for verifier caches. The private key stays with the signer.
After switching issuance, keep the old public key available until the last token signed with the old key has expired, including any allowed clock skew. Count from the last old-key issuance, not the start of deployment. Non-expiring signed artifacts need a separate replacement or retention policy.
The token held by the client still contains its original signature. Looking it up doesn’t rewrite that signature under the new key. Here, draining usually means expiration or replacement, rather than rehashing a database row. Auth0’s signing-key rotation provides a concrete example of publishing current and next verification keys before the switch.
Encryption keys: migrate the stored data
With reversible encryption, the old key can recover the plaintext. A batch job can read each encrypted record, decrypt it with the old key, and encrypt it with the new one without waiting for a client request. It needs access to the old key and all decryption metadata, such as the nonce, authentication tag, and any associated data. Google Cloud KMS documents this re-encryption workflow.
Use the encryption library’s rules for fresh nonces and authenticated encryption. Commit the new ciphertext and its key-version metadata together, and use a conditional update or locking so the job doesn’t overwrite a concurrent application write. Make the job resumable and count what still depends on the old version.
Creating a new key version doesn’t perform that migration automatically. Cloud KMS explicitly separates rotation from re-encryption. Keep the old decryption key available while old ciphertext or backups still need it. Destroying that key prematurely can make the data unrecoverable.
With envelope encryption, a data encryption key encrypts the payload, and a key encryption key wraps that data key. Rotating the wrapping key can involve rewrapping data keys while leaving the payload unchanged. Rotating the data key itself requires re-encrypting the payload. Be clear about which layer you’re retiring.
That is the difference from our HMAC case: ciphertext and its old key can recover the original input; a token fingerprint and its HMAC key cannot. Hash-only tokens must be presented again, replaced, or allowed to expire.
Service credentials: move callers and check new connections
For a database password or a third-party API credential, the receiving service controls what is accepted. If it supports two active credentials, provision the replacement first, give it the required permissions, and test it before updating callers.
Then update every caller, including workers and infrequent jobs. Refresh cached secrets and deliberately test fresh connections. A healthy connection pool can hide stale credentials because an established connection may remain usable after its original password stops working.
AWS Secrets Manager’s database rotation strategies illustrate the distinction: changing one user’s password can introduce a brief mismatch between the database and callers, while alternating between two users allows overlapping valid credentials. That overlap also requires keeping their permissions consistent.
Once callers use the replacement, handle old connections or sessions according to the provider’s behavior, then revoke the old credential. Check whether revocation also terminates existing sessions; don’t assume it does. There is no token-fingerprint migration in this case.
Rollback must preserve work already produced
Before switching, decide whether the rollback version can consume work created with the replacement. An older binary may understand the old secret but not the new key ID or storage format.
In the HMAC example, after K2 has been used, rolling back to a configuration that only knows K1 will break the new or migrated credentials.
You can switch the writer back to K1 while retaining both accepted keys. If you do, pause migration or explicitly change its target; don’t leave different processes migrating in opposite directions. Any rollback binary must also understand the prefixed storage format.
Backups need the same thought. Restoring an older database can bring back fingerprints under a retired key. Decide whether recovery will provide access to that key or require affected credentials to be replaced.
For encryption, a rollback reader may need to decrypt new-key ciphertext. For signing, it may need to verify new-key tokens. For service credentials, rolling back must not reload a credential the provider has already revoked.
A leaked key changes the decision
This rollout is for routine rotation, where keeping the application working is the goal. If a secret was exposed, keeping it accepted may conflict with the incident’s containment plan.
For this fingerprinting scheme, stealing the HMAC key alone doesn’t let someone mint arbitrary accepted tokens: they would still need a matching database record. With leaked fingerprints, though, the key lets them check token guesses offline. High-entropy tokens make guessing impractical; short codes don’t have that protection.
If the application itself was compromised, assume the investigation may also involve incoming tokens, database access, and workload credentials. Remove the attacker’s access before distributing replacement secrets. Token revocation or replacement may be necessary even if it interrupts clients. OWASP’s secrets-management guidance discusses containment and exposed-secret removal.
Decide what proves the rotation is finished
Track the active version and loaded key set per process, new work still using the old version, remaining dependencies, migration failures, and authentication or decryption failures. Keep tokens, plaintext, and key bytes out of logs.
Before a rollout, I’d test the transitions directly:
- Create work with the old version and consume it after the switch.
- Create work with the new version and consume it on every configuration that can run during the rollout.
- Race migrations with updates or revocation. In the HMAC case, check that unknown, expired, and revoked tokens still fail under every accepted key.
- Exercise the intended rollback against work already created with the replacement.
- Remove the old version in a test environment and verify both the surviving dependencies and those deliberately invalidated.
The shared sequence is expand, switch, drain, retire. The retirement condition must be specific: no usable old-key fingerprints, no unexpired old-key tokens, no required old-key ciphertext, or no callers and sessions depending on the old credential. A successful deployment alone doesn’t establish any of those.
The examples and rollout advice here describe a general design. Adapt them to your application’s storage model and deployment process.