← Seth Bergman

Operations · Reading the signal

no rekey configuration found — because the rekey had finished

The error handler deleted the only keys that still worked.

The error that meant it worked

Four times across two releases, Vault and AWS reported exactly what had happened and I read the report backwards. None of these were bugs in those systems.

While writing a script to rotate Vault's recovery keys, I destroyed a set of recovery keys.

Local cluster, one I tear down several times an hour, so nothing was actually lost. But the script did what it was written to do, the error handling fired correctly, and the keys were gone anyway.

Recovery keys are the five secrets that, three together, let you mint a root token when every other way in is gone. Re-issuing them is a one-way door: the instant the operation finishes, the old five stop working. If you did not capture the new ones, nobody administers that cluster again.

Vault has a second phase for exactly this. Ask for verification and the new shares are issued but held back — they do not take effect until you hand a quorum of them in. Mistype one and nothing happens; the old shares are still live and you get to try again.

You cannot ask for it from the CLI. vault operator rekey has a -verify flag for the handback and nothing to request verification when you start, because require_verification only exists in the API. That is most of why I was scripting the thing rather than typing it.

So the script fed in the old shares, got the new ones back, and started submitting them for verification, watching for the JSON field that says you are done. It never appeared. Here is what Vault returns on the share that completes verification, when you have asked for -format=json:

Rekey verification successful. The rekey operation is complete and the
new keys are now active.

English. No JSON, no complete field, nothing to parse. The loop shrugged and submitted the next share:

Code: 400. Errors:

* no rekey configuration found

There was no rekey configuration because the rekey had finished.

My error handler saw a 400, decided verification had failed, printed a reassuring note that the old shares were still the live ones, and threw the new set away. Every part of that was wrong. The old shares were already dead. The new ones were the only keys that would ever open that cluster again, and the failure path deleted them.

Ten minutes to fix — write the new shares to disk before verification rather than after, and treat the English sentence as a completion signal alongside the JSON one. What stuck with me was that the 400 was correct. Vault answered the question I asked. There genuinely was no rekey configuration. I had just decided in advance what its absence would mean.

Every one of these was a true statement that I understood as its opposite.

S3 · object lockAn absence that means someone was here

The same repository ships audit-trail anchors to S3 under Object Lock in COMPLIANCE mode. Nothing can shorten or lift that retention — not the operator, not the account root, not somebody holding the same key the shipper runs with. The point is that the record outlives whoever wants it gone.

Before building on it I wrote a throwaway script to check that the emulator I test against actually enforced any of this, rather than accepting the configuration and shrugging. That is the only reason I found the next part.

AttackObject LockWhat stops it
delete versionRefusedCOMPLIANCE retention
overwrite keyPermittedReading the version written first
delete, no versionPermittedAn IAM deny, and listing versions

The third row is the interesting one. A delete without a version id does not remove anything — it writes a delete marker over the key, and the data sits underneath untouched. Object Lock has no reason to refuse that, so it does not.

The catch is that a marked key disappears from list-objects-v2 and 404s on head-object. Those are the two calls you would naturally build a fetch tool on. They are the two I had built mine on. So the credential that ships the evidence can hide all of it without deleting a byte, and the tool you reach for during an incident tells you there is nothing there.

What it said

no anchors found

What was true

intact, one layer down

and somebody had just tried to delete them

You would like those two situations to look different. They do not. The tool lists object versions now, reads the anchors out from under the markers, and makes noise about the markers, because a delete attempt on an audit trail is the most interesting thing that will happen all week.

Vault · healthHealthy, and failing every request

A Vault node that has lost quorum is still unsealed and still listening. It just cannot agree with anyone about anything, so real requests come back 500.

Ask it sys/health?standbyok=true and you get a 200. The AWS target group matches 200,429, since 429 is how a healthy standby announces itself. So the node stays in the pool: healthy by every signal the infrastructure collects, serving errors to everything that arrives.

Again, nothing lied. standbyok=true means "a standby is fine", the node is arguably a standby, and 200 is a defensible answer. The health check is not measuring what its name suggests it measures. What catches this is an alert on sum(vault_core_active) < 1, which counts leaders instead of asking nodes how they feel.

Raft · autopilotTwo of three healthy, and no quorum

Vault ships Raft autopilot with cleanup_dead_servers = false, and honestly that is the right default. On fixed machines a node that vanishes is usually coming back as itself, and evicting it automatically causes more problems than it solves.

Both cloud profiles in this repository set the Raft node id to something the machine does not keep — the EC2 instance id, the scale set VM name. Every replacement therefore arrives as a brand new voter, and the one it replaced sits in the configuration forever, dead and still counted.

Then you run an instance refresh, the most ordinary thing an operator does all month:

StepVotersQuorum needsLive
startA B C23ok
terminate AA̶ B C22no margin
launch A2A̶ B C A233no margin
terminate BA̶ B̶ C A232quorum lost

It goes down partway through the second node, while the console still shows two of three healthy and the rollout progressing normally. min_healthy_percentage = 67 cannot help, because the scaling group is counting instances it can see and Raft is counting voters it cannot. Both numbers are accurate. They are counting different things, and only one of them decides whether the cluster serves.

I found this on a laptop, incidentally. No cloud account and no spend — a local three-node cluster, killing a container, and watching list-peers still show the dead node as a voter a minute later.

There is a footnote I am slightly embarrassed by. I wrote the fix up before testing it and got the sequence wrong. The replacement actually joins as a non-voter, which satisfies the quorum floor on its own; the dead voter is pruned; the replacement is promoted after that. The voter count never reaches four. I only know because the first assertion I wrote encoded my explanation and then failed against a cluster that was working perfectly.

What I do differently now

These are not test failures. My tests were fine, and in three of the four cases there was no test to be wrong — the mistake was in the model I had of what a correct answer implied. So the habits are different from the ones I would reach for after a bad assertion.

Ask what else produces this response

Before writing an error branch, work out the full set of conditions that reach it. A 400 from a verification endpoint has at least two causes: verification failed, and verification is over. If a branch cannot distinguish them, it must not act as though it can — and the destructive action in particular belongs somewhere that a wrong guess cannot reach.

Prefer signals with one meaning

Count leaders, not health checks. List versions, not objects. Every one of these bugs came from a signal that was cheap to obtain and ambiguous by nature, and the fix in every case was a slightly more expensive signal that could only mean one thing.

Make the irreversible step last

The rekey was survivable the moment the new shares were written to disk before verification instead of after it. Nothing about the misread 400 changed. It just stopped mattering, because by then the only thing the failure path could destroy was a file I already had a copy of.

That last one is the only general lesson I actually trust from this. You will misread something eventually. Arrange the sequence so that when you do, the thing you cannot get back has already been written down.