Operations · Reading the signal
no rekey configuration found — because the rekey had finished
The error handler deleted the only keys that still worked.
The error that meant it worked
Four times across two releases, Vault and AWS reported exactly what had happened and I read the report backwards. None of these were bugs in those systems.
While writing a script to rotate Vault's recovery keys, I destroyed a set of recovery keys.
Local cluster, one I tear down several times an hour, so nothing was actually lost. But the script did what it was written to do, the error handling fired correctly, and the keys were gone anyway.
Recovery keys are the five secrets that, three together, let you mint a root token when every other way in is gone. Re-issuing them is a one-way door: the instant the operation finishes, the old five stop working. If you did not capture the new ones, nobody administers that cluster again.
Vault has a second phase for exactly this. Ask for verification and the new shares are issued but held back — they do not take effect until you hand a quorum of them in. Mistype one and nothing happens; the old shares are still live and you get to try again.
You cannot ask for it from the CLI. vault operator rekey
has a -verify flag for the handback and nothing to request
verification when you start, because require_verification
only exists in the API. That is most of why I was scripting the thing
rather than typing it.
So the script fed in the old shares, got the new ones back, and
started submitting them for verification, watching for the JSON field
that says you are done. It never appeared. Here is what Vault returns
on the share that completes verification, when you have asked for
-format=json:
Rekey verification successful. The rekey operation is complete and the
new keys are now active.
English. No JSON, no complete field, nothing to parse.
The loop shrugged and submitted the next share:
Code: 400. Errors:
* no rekey configuration found
There was no rekey configuration because the rekey had finished.
My error handler saw a 400, decided verification had failed, printed a reassuring note that the old shares were still the live ones, and threw the new set away. Every part of that was wrong. The old shares were already dead. The new ones were the only keys that would ever open that cluster again, and the failure path deleted them.
Ten minutes to fix — write the new shares to disk before verification rather than after, and treat the English sentence as a completion signal alongside the JSON one. What stuck with me was that the 400 was correct. Vault answered the question I asked. There genuinely was no rekey configuration. I had just decided in advance what its absence would mean.
Every one of these was a true statement that I understood as its opposite.
S3 · object lockAn absence that means someone was here
The same repository ships audit-trail anchors to S3 under Object
Lock in COMPLIANCE mode. Nothing can shorten or lift that
retention — not the operator, not the account root, not somebody
holding the same key the shipper runs with. The point is that the
record outlives whoever wants it gone.
Before building on it I wrote a throwaway script to check that the emulator I test against actually enforced any of this, rather than accepting the configuration and shrugging. That is the only reason I found the next part.
| Attack | Object Lock | What stops it |
|---|---|---|
| delete version | Refused | COMPLIANCE retention |
| overwrite key | Permitted | Reading the version written first |
| delete, no version | Permitted | An IAM deny, and listing versions |
The third row is the interesting one. A delete without a version id does not remove anything — it writes a delete marker over the key, and the data sits underneath untouched. Object Lock has no reason to refuse that, so it does not.
The catch is that a marked key disappears from
list-objects-v2 and 404s on head-object.
Those are the two calls you would naturally build a fetch tool on. They
are the two I had built mine on. So the credential that ships the
evidence can hide all of it without deleting a byte, and the tool you
reach for during an incident tells you there is nothing there.
What it said
no anchors found
What was true
intact, one layer down
and somebody had just tried to delete them
You would like those two situations to look different. They do not. The tool lists object versions now, reads the anchors out from under the markers, and makes noise about the markers, because a delete attempt on an audit trail is the most interesting thing that will happen all week.
Vault · healthHealthy, and failing every request
A Vault node that has lost quorum is still unsealed and still listening. It just cannot agree with anyone about anything, so real requests come back 500.
Ask it sys/health?standbyok=true and you get a
200. The AWS target group matches
200,429, since 429 is how a healthy standby announces
itself. So the node stays in the pool: healthy by every signal the
infrastructure collects, serving errors to everything that arrives.
Again, nothing lied. standbyok=true means "a standby is
fine", the node is arguably a standby, and 200 is a defensible answer.
The health check is not measuring what its name suggests it measures.
What catches this is an alert on sum(vault_core_active) < 1,
which counts leaders instead of asking nodes how they feel.
Raft · autopilotTwo of three healthy, and no quorum
Vault ships Raft autopilot with
cleanup_dead_servers = false, and honestly that is the
right default. On fixed machines a node that vanishes is usually coming
back as itself, and evicting it automatically causes more problems than
it solves.
Both cloud profiles in this repository set the Raft node id to something the machine does not keep — the EC2 instance id, the scale set VM name. Every replacement therefore arrives as a brand new voter, and the one it replaced sits in the configuration forever, dead and still counted.
Then you run an instance refresh, the most ordinary thing an operator does all month:
| Step | Voters | Quorum needs | Live | |
|---|---|---|---|---|
| start | A B C | 2 | 3 | ok |
| terminate A | A̶ B C | 2 | 2 | no margin |
| launch A2 | A̶ B C A2 | 3 | 3 | no margin |
| terminate B | A̶ B̶ C A2 | 3 | 2 | quorum lost |
It goes down partway through the second node, while the console
still shows two of three healthy and the rollout progressing normally.
min_healthy_percentage = 67 cannot help, because the
scaling group is counting instances it can see and Raft is counting
voters it cannot. Both numbers are accurate. They are counting
different things, and only one of them decides whether the cluster
serves.
I found this on a laptop, incidentally. No cloud account and no
spend — a local three-node cluster, killing a container, and watching
list-peers still show the dead node as a voter a minute
later.
There is a footnote I am slightly embarrassed by. I wrote the fix up before testing it and got the sequence wrong. The replacement actually joins as a non-voter, which satisfies the quorum floor on its own; the dead voter is pruned; the replacement is promoted after that. The voter count never reaches four. I only know because the first assertion I wrote encoded my explanation and then failed against a cluster that was working perfectly.
What I do differently now
These are not test failures. My tests were fine, and in three of the four cases there was no test to be wrong — the mistake was in the model I had of what a correct answer implied. So the habits are different from the ones I would reach for after a bad assertion.
Ask what else produces this response
Before writing an error branch, work out the full set of conditions that reach it. A 400 from a verification endpoint has at least two causes: verification failed, and verification is over. If a branch cannot distinguish them, it must not act as though it can — and the destructive action in particular belongs somewhere that a wrong guess cannot reach.
Prefer signals with one meaning
Count leaders, not health checks. List versions, not objects. Every one of these bugs came from a signal that was cheap to obtain and ambiguous by nature, and the fix in every case was a slightly more expensive signal that could only mean one thing.
Make the irreversible step last
The rekey was survivable the moment the new shares were written to disk before verification instead of after it. Nothing about the misread 400 changed. It just stopped mattering, because by then the only thing the failure path could destroy was a file I already had a copy of.
That last one is the only general lesson I actually trust from this. You will misread something eventually. Arrange the sequence so that when you do, the thing you cannot get back has already been written down.