Testing · Terraform
15 + 1 + 8 <= 24
An assertion with no input that could make it fail.
Nobody had watched these tests fail
A mutation table is a list of claims about your test suite: break this, and that run goes red. Mine had thirteen rows, and a heading admitting none of them had ever been run. Six assertions did not survive the first time anyone checked.
The repository this comes from has one rule about what "done" means: a feature is done when there is a test that fails if it breaks. Not when the code exists. To keep that honest, each Terraform suite carries a table — deliberate break in one column, the test run that catches it in the other.
The AWS table was built the right way round: every row was added after watching that run go red. The Azure table was not, and its own heading said so:
The mutation table is not verified yet. An assertion that has never been watched to fail is an assertion nobody has shown to work.
I wrote that sentence. Then I shipped a release with it still true.
Verifying it needs no credentials, no cluster and no money — the suite is mocked providers and it finishes in about two seconds. So I ran the thirteen rows, plus one mutation for every run the table had never listed: 36 in all, each applied, run, and reverted. Thirty-three landed on the assertion that claimed them. Three were refused by the provider before any assertion ran. Six assertions stayed green through the break they named.
The one that could not fail
Azure Key Vault names are capped at 24 characters and are globally unique, so the module builds one from a truncated cluster name plus eight hex characters of randomness. The test looked like this:
assert {
condition = length(substr(replace(lower(var.cluster_name), "/[^a-z0-9-]/", ""), 0, 15)) + 1 + 8 <= 24
error_message = "Key Vault name would exceed Azure's 24-character limit: 15-char prefix + dash + 8 hex is the budget."
}
Read it as arithmetic rather than as a test. substr(x, 0, 15)
returns at most fifteen characters. Fifteen, plus one for the dash, plus
eight for the suffix, is 24. The condition is 24 <= 24 for
every input in existence.
There was no cluster name, no configuration and no future edit that could turn that assertion red. It had been green since the day it was written, and it would have stayed green forever.
The mutation that proves it: widen the module's real budget from 15 to 20 — the change someone makes the first time a customer wants a longer cluster name — and every run still passes, while the module now generates a 29-character name that Azure rejects at apply. Which is precisely the failure this assertion was written to prevent, and it had been documented as covered.
The flaw is not the arithmetic. It is that the test re-derived the module's naming rule instead of reading it. A test that recomputes what the code computes is testing the recomputation.
The fix is to make the value legible to the test. The module now exposes the prefix as a local, and the assertion reads it:
variables {
# At the 15-character default, every budget yields the same
# 15-character prefix. Only a name at the cap can tell them apart.
cluster_name = "vault-reference-platform-azure"
}
assert {
condition = length(local.key_vault_name_prefix) + 1 + 8 <= 24
error_message = "Key Vault name would exceed Azure's 24-character limit: 15-char prefix + dash + 8 hex is the budget."
}
That comment is the second half of the lesson. Reading the local is not enough on its own: the default cluster name is exactly fifteen characters, so truncating at 15 and truncating at 20 produce identical output. The assertion needed an input long enough for the two to disagree. Without it, the fixed test would have been just as green as the broken one, for a subtler reason.
The one the provider was already enforcing
Next to it, guarding the key that every snapshot is encrypted under:
condition = azurerm_key_vault.vault_autounseal.soft_delete_retention_days >= 7
The module sets 90. Shortening that to a week is a real regression: it is the difference between recovering a key someone deleted last month and not. So I mutated it to 3 and expected red.
I got red, from the wrong place:
Error: expected soft_delete_retention_days to be in the range (7 - 90), got 3
The azurerm provider validates that field's range in its own schema — which mocked providers still run, because mocking replaces what a provider does, not what it will accept. So the assertion sat exactly on a bound something else already enforced. Every value the provider allows passes it. Every value that would fail it never reaches it.
It now asserts >= 30, above the floor, where shortening
90 days to the minimum does break it.
The same trap, one file over. A catch-all deny rule has to sit below every allow in an Azure NSG, and the table's row read "the deny rule moved above the allows". Azure's priority floor is 100 and the API allow holds it, so above every allow is not a configuration that exists — the provider refuses to create it. The mistake a contributor can actually make is a deny at 105: the API stays reachable, so nothing looks wrong, while Raft at 110 and the health probe at 120 are both denied. A cluster that never forms, behind a load balancer that ejects every node.
The assertion compared the deny rule against one allow. 105 walked straight through it.
The ones watching the wrong object
Two more, both of the same shape — the assertion looked at an input rather than at the thing built from it.
"The Vault API is not reachable from the whole internet" was asserted
as !contains(var.allowed_cidr_blocks, "0.0.0.0/0"): a claim
about a variable's default. Replace the rule's
source_address_prefixes = var.allowed_cidr_blocks with
source_address_prefix = "Internet" and the API is open to
everyone while the default stays dutifully RFC1918. Green.
And Raft discovery. Azure's peer discovery enumerates a scale set, which go-discover expresses as a resource group and a scale set name — either alone is rejected. The run asserted that the rendered cloud-init named the scale set, and that it carried a subscription. It never asserted the resource group. Dropping it leaves a line that still reads like scale-set discovery, that Vault starts happily with, and that forms no cluster.
The gap that had outlived its reason
The sixth was not a wrong assertion but a missing one, and it came with a comment explaining why:
# Not asserted here: that health_probe_id points at the Vault probe
# specifically. That comparison needs the probe's computed ID, and
# resolving it means apply mode, which this module does not survive
# under mocks.
That was true when it was written. It had stopped being true two commits later, when a fix to the storage account's identity mock unblocked apply mode — which two other runs in the same file had been using ever since.
Meanwhile, deleting the health_probe_id line entirely left
all 24 runs green. Without a probe attached, Azure's instance repair falls
back to asking whether the VM is running, and a Vault node that is up and
sealed is exactly what this cluster's repair exists to notice.
Comments explaining why something is not covered age badly. Nothing fails when the reason expires.
A note on mocked providers
Three of the 36 mutations were refused by the provider's schema before
any assertion evaluated, and the way that surfaces is worth knowing: a
schema error fails the first run in each file and skips
everything after it. Break the repair grace period in
compute.tf and the report says
scale_set_is_pinned_and_does_not_autoscale and
the_autounseal_key_cannot_be_purged failed — two runs that
have nothing to do with it.
The error text names the real line. The run names do not. It is a five-minute confusion the first time and a thirty-second one after you have written it down.
What these have in common
| The claim | Why it held anyway |
|---|---|
| the name fits in 24 characters | the test re-derived the rule; substr guarantees it |
| retention is at least 7 days | the provider refuses anything under 7 |
| the deny rule sits below the allows | compared against one allow of three |
| the API is not open to the internet | read the variable, not the rule |
| discovery uses scale-set mode | half the selector was never asserted |
| the probe is wired to repair | documented as uncovered, reason since expired |
Not one of these is a lazy test. Every one names the right property, with a comment explaining why the property matters, and every comment is correct. They fail in a narrower way: they have no input that makes them red.
That is a different failure from the one I wrote about last time, where an assertion made a claim that was wrong. These claims are all true. They are just not being checked, and nothing about reading them says so. The only way to find out is to try.
The tells
The test recomputes what the code computes. If the assertion contains the same expression as the module, it will agree with the module by construction. Expose the value — a local, an output — and read it.
The bound belongs to somebody else. Provider schemas, platform floors, variable validations: all of them reject the bad value before your assertion sees it. An assertion has to sit strictly inside the enforced range to be doing any work.
The default is the fixture. An assertion exercised only at the default input can be blind to a whole class of change — the 15-character cluster name that made two different truncations indistinguishable. Ask what input would separate correct from broken, and pass that one.
The comment says why it is not covered. That reason was true once. Nothing re-checks it, and nothing goes red when it stops being true.
And the table itself
The uncomfortable part is not any of the six. It is that verifying the whole table cost about an hour, needed no credentials and no cloud account, and the suite it exercises runs in two seconds. Nothing stood between the unverified table and the verified one except doing it.
A mutation table nobody has run is worse than no table at all, because it reads as evidence. Mine sat in a file whose heading said, accurately, that it was not.
The table now records what each mutation actually did, including the
three the provider rejects and the row where the catching run turned out
to be a different one than claimed. It also records what is still not
covered: the pinning test's comment says there is deliberately no
autoscale rule in the module, and nothing enforces that — adding one
passes all 25 runs. terraform test asserts on what a
configuration contains, not on what it lacks. That is written down now
too, as an intention rather than a guarantee.
None of this changes the honest gap underneath: that profile has still never been applied to a real Azure subscription, and mocked providers create nothing. It changes how much the word "tested" is worth in the sentence that says so.