← Seth Bergman

Operations · Seal migration

Seal Type transit — and it will not unseal itself

The cluster reports the new seal before it commits to it.

Checking it worked is what broke it

Migrating a Vault cluster between seal types has a window a few seconds long in which every obvious way to confirm it worked will undo it.

I had just moved a three-node cluster from Shamir back to Transit auto-unseal. All three nodes reported transit. All three were unsealed. The data was there.

What I wanted to know was whether auto-unseal actually worked — as opposed to being configured, which is a different claim and the one that had just become true. So I did the obvious thing: restarted a node and watched to see whether it came back on its own.

It came back sealed.

That is the correct symptom of a broken seal stanza, an expired Transit token, or an unreachable unseal service. I checked all three. They were fine. The node was refusing to do the thing its configuration plainly told it to do, and the reason was in the log, which is the only place it appears:

[WARN] core: entering seal migration mode; Vault will not automatically
unseal even if using an autoseal

The migration had not finished. Not because anything failed — it finishes on its own, and quickly — but because I restarted a node during the seconds that takes, and a node restarted inside that window re-enters migration mode. In migration mode Vault will not auto-unseal, deliberately, even when it can.

The check produced the failure it was checking for.

And it produced it in the shape of a different and much worse problem: a node that looks like it has a broken seal configuration, when what it has is a correct one and bad timing.

The sequenceFour steps, three of them surprising

Changing seal type means telling every node about the new seal, restarting them, and unsealing each one with -migrate. Written out:

  1. Set the seal stanza on every node.
  2. Restart every node. They come up sealed, already reporting the new type.
  3. Unseal every node with vault operator unseal -migrate.
  4. Wait for a leader to finalise it.

I guessed wrong on three of those four before running it.

Do not stop the standbys

The instinct, for any operation this dangerous, is to reduce the number of moving parts: take the standbys down, do the delicate thing on one node, bring them back.

On three nodes that leaves one, which is not a quorum. No leader is elected, and finalising the migration is something a leader does. My first attempt produced a cluster that was unsealed, leaderless and permanently half-migrated — worse than any state I was trying to avoid.

Every node needs -migrate, not just the active one

Submitting an ordinary unseal key to a standby gets you:

Code: 500. Errors:

* migrate option not provided and seal migration is in progress

A 500, which reads as a server fault, for a node correctly reporting that it is in the middle of doing what you asked.

It is not over when the last node unseals

This is the one that makes the restart destructive. Every node is unsealed, every node reports the new seal type, and the migration is still in progress:

vault status

Seal Type transit Sealed false

sys/seal-status

"migration": true

both true, one of them the whole answer

The flag clears when a leader finalises the migration, which on a healthy cluster takes seconds. It is a small window. It is also exactly the window you are in at the moment you finish the last unseal and go looking for confirmation that it worked.

The fixPoll the state, don't infer it

The class of mistake is reading something that is true now and will be true differently in a moment, then acting on it. The seal type is that kind of state. So is "every node is unsealed".

sys/seal-status answers the actual question:

curl -s "$VAULT_ADDR/v1/sys/seal-status" | jq -r '.migration'

Wait for false before doing anything else — before restarting a node, before closing the maintenance window, before running the check that tells you it worked.

There is a second-order version of the same problem, because it bit me straight afterwards. Having written a script that migrates a cluster, I pointed it at one that was already half-migrated — the state my first attempt had produced — and it refused:

ERROR: already on transit; nothing to migrate

It read the seal type, saw the target, and concluded there was nothing to do. The nodes reported transit because they were halfway there, not because they had arrived. A tool that performs a dangerous operation should expect to meet clusters in the middle of that operation, since that is exactly when someone reaches for it. It resumes now rather than refusing.

PostscriptMy own contribution

The wait I wrote for that flag could not succeed. It read it like this:

jq -r '.migration // empty'

jq's // treats false the same as absent, so a finished migration returned an empty string and the loop never matched. It timed out after three minutes against a cluster that had finalised in about six seconds.

A wait that cannot terminate and a check that cannot fail are the same bug in different clothes, and I have now written both in the same repository.