Mutation Testing for Infrastructure Code — 153 Planted Bugs, Four Tests That Could Not Fail
By Vladimir Mikhalev · Solutions Architect · Docker Captain · CNCF Ambassador · IBM Champion
Fifteen of my public repositories carried a green CI badge and a README line saying they were verified on every push. Not one of them ran a step that could fail because the code did the wrong thing. Their workflows ran actionlint, tflint, shellcheck, hadolint and a Docker build. Lint is an opinion about text. The badge was being read as a statement about behaviour, and I was the one who had put it there. So I wrote the tests, then ran mutation testing on them by hand: 153 deliberate breaks, one per promise, each required to turn a test red. Four tests stayed green. Two shipped defects surfaced on the way, one of them almost five years old.
This is the same blindness I wrote about in health checks that cannot fail, one layer further out. A health check that cannot go red is a lie told to the scheduler. A test that cannot go red is a lie told to the reviewer, and the reviewer is usually you, six months later, reading a green run as proof.
What a test that cannot fail looks like
The log cleaner for IIS and Exchange makes three promises. Old logs of the named kinds go, everything else stays, and a WhatIf run touches nothing. The first run caught its WhatIf test in this shape, checking the files and nothing else:
# clean-server-logs.Tests.ps1, the shape the plant caught: the files, not the claimIt 'deletes nothing under -WhatIf' { & $Script -Path $Root -Days 30 -WhatIf $OldLog, $OldEtl, $OldBlg, $NewLog | ForEach-Object { Test-Path -LiteralPath $_ | Should -BeTrue }}The plant rewrote the script’s own ShouldProcess check so that it always said yes, and the script stopped asking before it deleted. The test stayed green. Remove-Item honours WhatIf by itself, so the files survived either way, and the test was proving a promise PowerShell keeps rather than one the script keeps.
# clean-server-logs.Tests.ps1: the count is the only place a bypassed guard showsIt 'deletes nothing under -WhatIf, and does not claim to have' { $out = & $Script -Path $Root -Days 30 -WhatIf $OldLog, $OldEtl, $OldBlg, $NewLog | ForEach-Object { Test-Path -LiteralPath $_ | Should -BeTrue } $out | Should -Match '^Deleted 0 file\(s\)'}With the guard bypassed, the script counts every file it believes it deleted, and the last line goes red. In banking, telecom and hyperscaler support I have watched the same thing happen with every generation of tooling: the instrument is trusted because it was installed, not because it was ever seen to fire.
The difference was one assertion.
Why this keeps happening
Coverage is the number teams report, and it measures the wrong thing. In what was then the largest study of the question, Inozemtseva and Holmes generated 31,000 test suites for five Java systems of up to 724,000 lines, and used mutation testing to measure how well each suite caught faults. Once the size of the suite was held constant, coverage correlated with that ability only weakly to moderately, and stronger forms of coverage added nothing. Their conclusion was that coverage should not be used as a quality target. A test that runs a line and asserts nothing about it still covers the line.
Infrastructure code gets even less. Mutation testing is old and well understood in application code. PIT does it for Java, Stryker for JavaScript and TypeScript, mutmut for Python. A Terraform module, a PowerShell maintenance script and a CI example meant to be copied have no tool that understands what matters in them, and a random mutation is the wrong instrument anyway. Reversing a comparison somewhere in a module tells you little. Making a database reachable from the internet tells you at once whether the test guarding that promise works.
The badge hides the rest. A workflow that only lints goes green for the same reasons as one that tests, and a reader cannot tell which it was. The OpenSSF Best Practices questionnaire asks eight questions about an automated test suite. Answered honestly for 89 of my repositories on 25 September, the count was 74 with a suite and fifteen without.
A green badge on a repository with no test says lint. It does not say verified.
Risk and blast radius
Direct exposure is the change record. The Active Directory script in this set disables dormant accounts, and its README tells the reader to run it with WhatIf first. In a regulated environment, that dry-run report is the artefact a change advisory board signs before anything is disabled. SOC 2 CC8.1 asks whether changes are tested before they are implemented, and ISO 27001 control 8.29 asks for security testing in development and acceptance. A green run and a test file satisfy both. Neither asks whether the test could have gone red.
Systemic exposure is the copy. These repositories are templates. A Terraform pipeline or an OS update job is meant to be copied into somebody else’s account and CI, and the copy inherits the tests along with the code. A guard that no test can see ships into every copy as decoration, and the person copying it has no way to know which guards those are. The OS update job is the sharpest case. It logs into a host over SSH and upgrades its packages, so the setting that decides whether it will talk to an impostor matters more than anything else in it, and nothing tested it.
A control nobody has seen fail is a rumour with a badge.
What mutation testing found
On the first runs five planted violations went unnoticed, and behind them were four tests that could not fail. All 153 are caught now, and so are the six added since. Each finding below is on the ledger with the rule it produced, and all fifteen repositories now pass the OpenSSF badge. The evidence page lists 90 of 90 registered repositories as passing, every answer public and each one naming the file that makes it true.
Terraform: testing a plan with no AWS account
The eight Terraform pipelines are tested with terraform test, which since Terraform 1.7, released on 17 January 2024, can replace a provider with a mock. Each test plans the configuration with its default variables, against mocked providers, and asserts what would be built. No account, no credentials, nothing created:
# tests/posture.tftest.hcl in the RDS pipeline, shortened: mocked providers, no AWS accountmock_provider "aws" {}
run "the_defaults_are_the_secure_ones" { command = plan
assert { condition = aws_db_instance.db_instance_1.storage_encrypted == true error_message = "db_instance_1 stores data unencrypted" }
assert { condition = aws_db_instance.db_instance_1.publicly_accessible == false error_message = "db_instance_1 is reachable from the internet" }}Between them the eight files make 164 assertions: KMS key rotation, S3 public access blocks, encryption and versioning, the state lock’s encryption and point-in-time recovery, IMDSv2 on instances, invalid-header dropping on load balancers, TLS policy on listeners, flow logs, EKS secrets encryption. On the first run 109 plants broke the promises behind them, and every one was caught.
The assertions were generated from each configuration’s own resources, which is how the more useful output appeared. Generated that way, they could only assert what the defaults deliver, so every place those defaults fall short became a line in the README instead of a silent gap. SSH is open to the whole internet in four EC2 pipelines. Redis is not encrypted at rest or in transit in two. The EKS API endpoint is public. None of these is hidden any more, and none is claimed.
The load balancers show why a list of plants is never finished. Version 1.2.0 moved three of them to AWS’s TLS 1.3 policy from 2021 and called it the policy AWS recommends. It was not. AWS now recommends its hybrid post-quantum policies, and version 1.3.0 moved the three to one of them. The assertion now requires both TLS 1.3 and a post-quantum key exchange in the policy’s name, and two new plants in each pipeline try the old policy and a TLS 1.2 one. Six more plants, 159 in all, and a changelog entry that says the previous release was wrong.
# 17-alb.tf in the GitLab pipeline, inside the HTTPS listenerssl_policy = "ELBSecurityPolicy-TLS13-1-2-Res-PQ-2025-09"A test that asserts only the good news is a brochure.
The dry run that reported a file it never wrote
The Active Directory script disables accounts that have not signed in for a set number of days, stamps them, and moves them to a holding OU. It supports WhatIf properly, through SupportsShouldProcess. Under WhatIf it printed how many dormant accounts it had written to its report.
It wrote no report.
WhatIf sets a preference variable for the whole script scope, and every cmdlet that supports ShouldProcess inherits it, including New-Item and Export-Csv. The account changes were correctly skipped. So was the report, the one thing a dry run exists to produce. The line after it announced a file that did not exist. A two-parameter fix exempts those two cmdlets from the dry run:
# disable-inactive-users-active-directory.ps1: the report is written even under -WhatIfNew-Item -Path $LogFolder -ItemType Directory -Force -WhatIf:$false | Out-Null$report | Export-Csv -LiteralPath $logFile -NoTypeInformation -Encoding UTF8 -WhatIf:$falseThe Pester suite that found it runs the script against a stand-in ActiveDirectory module that records every call instead of making it. That is what makes an AD script testable on a CI runner with no domain. Its tests assert the sequence of changes the script would make, in order, and that WhatIf makes none of them. For three weeks, since the script learned WhatIf on 3 September, the dry run had been a sentence with no file behind it.
The other defect was older. In the terminal tree, the Ctrl-C handler called tput reset, which wipes the scrollback and, on a terminal that does not answer it, stalls long enough for the script to be killed before the cursor comes back. It had been there since the first commit, in December 2021, and a headless test that runs the tree under util-linux script found it on its first run. The handler now restores the colours and the cursor, clears the screen, and exits with the signal’s code:
# terminal-christmas-tree.sh: give the terminal back on Ctrl-Crestore() { tput sgr0; tput cnorm; clear; }trap 'restore; exit 130' INTtrap 'restore; exit 143' TERMBoth had shipped. Both surfaced the first time a test asked what the script does, not what it prints.
Four tests that could not fail
The plants did their work on the tests, not the code. Five of them missed on the first runs, from four tests, and each of those tests looked right and proved nothing.
| Test | What the plant broke | Why the test stayed green | What it asserts now |
|---|---|---|---|
Log cleaner, -WhatIf | The script’s own ShouldProcess guard | The delete cmdlet honours the dry run on its own | The count the script reports |
| Mattermost retention | A code path the test never reached | The test never ran that path | A 20-day-old post that a 30-day retention keeps |
| Terminal tree size | The size of the tree | The check read the drawn output as a whole | The widest row drawn between cursor moves |
| OS update over SSH | The empty-scan guard before the login | Strict host key checking refused the host anyway | An impostor takes the target’s name, and the job must refuse |
The first is the WhatIf test above. In the Mattermost retention test, one plant broke a code path the test never reached. It was replaced by a plant the test could see, and by a 20-day-old post that a 30-day retention has to keep, because a retention job that deletes too much is the expensive failure, not the one that deletes too little. The terminal tree’s size check read the drawn output as a whole, and a plant that shrank the tree passed it. It now measures the widest row drawn between cursor moves.
The fourth is worth the whole exercise.
The host key check that sat behind a second lock
Two example jobs ship with the OS update pipeline, one for GitHub Actions and one for GitLab CI, and both SSH into a host and run apt-get upgrade. Each scans the host key, refuses to continue when the scan comes back empty, and then connects with strict host key checking on. The first versions, years ago, connected with it off.
# .github/workflows/00-os-update.yml.example, shortened: the scan, the guard, the lockssh-keyscan -H "$EC2_HOST" > ~/.ssh/known_hosts 2>/dev/nulltest -s ~/.ssh/known_hosts || { echo "::error::no host key from $EC2_HOST"; exit 1; }ssh -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=yes -o BatchMode=yes \ "$SSH_USER@$EC2_HOST" 'sudo apt-get update && sudo DEBIAN_FRONTEND=noninteractive apt-get upgrade -y'Two plants removed the empty-scan guard, one in each example. The test stayed green.
With the guard gone, strict host key checking refuses the unknown host anyway, so the job still stopped and nothing a test could see had changed. The guard was a second lock behind the first. Testing it proved nothing, because the thing that mattered was the first lock, and nothing tested that.
Strict host key checking exists for one case. The host that answers is not the host whose key you scanned. So the test now builds that case. The runner scans the real target’s key, then the target’s network name moves to an impostor container with freshly generated host keys and the same account, and only then does the job’s SSH step run:
# The target's name moves to the impostor between the key scan and the login.swap_to_impostor() { docker network disconnect "$RUN" "$RUN-target" >/dev/null docker network connect --alias target "$RUN" "$RUN-impostor" >/dev/null}The job has to fail, and nothing may run on the impostor. Now the plant that matters turns strict host key checking off. With that plant in place the job logs straight into the impostor and upgrades it. The test goes red, and the run reports the plant as caught. That is a man-in-the-middle test running on a free GitHub runner inside that job, for a pipeline that is copied into other people’s CI.
A security control is tested when an attacker’s move is shown to fail against it. Removing the control’s neighbour is not that.
Framework: promise, plant, and a run that has to go red
The fleet’s own engine has worked this way since 7 September, when a health check that could never fail made it a rule that every check ships with the test that plants its violation. There are about four hundred of them, now public. The fifteen repositories were where the rule had not reached. Any repository can take the same discipline in three steps.
Layer 1: write the promises and their plants (week 1)
Write down what the repository promises, one sentence per promise, in the terms a reader would use. “The database is not reachable from the internet”, not the name of the setting that makes it so. For each promise, write the smallest edit that breaks it. That edit is a plant. In a large organisation, start with the promises whose failure would be an audit finding rather than a bug: who can reach the data, what gets deleted, what a dry run records, and which host a job will log into. Every tested repository here keeps them in one tab-separated file beside its tests. Each line names a file, a piece of text in it, what to put there instead, and a sentence saying what the test must notice. This line from the RDS pipeline, shortened here to the lines that change, flips the default for public accessibility:
# tests/plants.tsv: file, old text, new text, what the tests must notice"00-variables.tf" "type = bool\n default = false" "type = bool\n default = true" "publicly_accessible flipped in the default of rds_publicly_accessible"The fields are JSON strings separated by tabs, so a plant can span lines and carry quotes without an escaping scheme of its own.
Owner: the repository’s maintainer.
Layer 2: break each promise on its own copy (week 2)
The runner that reads the plants is about eighty lines of Python. Script repositories share one, and the Terraform pipelines use a variant that runs terraform test in a pinned container. Its core runs the suite on an untouched copy, then applies each plant to a fresh copy and counts what stayed green:
# tests/plant-violations.py: the untouched copy first, then one plant per copywork = copy()r = run(cmd, work)if r.returncode != 0: sys.exit("the untouched code does not pass its own tests")
missed = stale = 0for fname, old, new, what in plants: work = copy() path = os.path.join(work, fname) text = open(path, encoding="utf-8").read() if old not in text: print(" STALE %s: the text this plant breaks is no longer in %s" % (what, fname)) stale += 1 else: open(path, "w", encoding="utf-8").write(text.replace(old, new, 1)) if run(cmd, work).returncode == 0: print(" MISSED %s" % what) missed += 1The untouched copy has to pass. Otherwise every plant “fails” for a reason that has nothing to do with the plant, and the run reports a perfect score over a broken suite. And every plant runs on its own copy, because two violations at once can mask each other.
Owner: platform engineer.
Layer 3: run the plants on every push, and fail on a stale one (week 3 onward)
A plant whose text is gone is stale, and stale fails the run. Code changes. A plant that silently stops applying is a test that silently stops being tested, which is the exact failure this whole exercise exists to catch. So the plants run in CI on every push, beside the tests:
# .github/workflows/verification.yml in the log cleaner# A test that cannot fail is not a test: each promise is broken on a# copy of the script, and the run fails if the tests stay green.- name: The tests, shown each violation they exist for run: python3 tests/plant-violations.py -- pwsh -NoProfile -NonInteractive -File tests/run.ps1When a plant survives, fix the test first. The survivor is information, not noise.
Owner: platform engineer, with security owning the plants that guard access: public endpoints, TLS policy and SSH host keys.
Tradeoffs
The plant runs are not free. Each plant rebuilds or replans on a fresh copy, so a repository with eighteen plants runs its suite nineteen times. On GitHub-hosted runners the slowest here, the OS update pipeline with two container images, spends about three minutes per push on its test and plants together. Public repositories pay nothing for Actions minutes. A private monorepo would want the plants on a schedule or on changes to the tested files, not on every push.
Skipping them costs what this post already showed: a dry run that could have been signed for three weeks with no report behind it, and an SSH guard whose only test proved nothing.
Three minutes a push, against a dry run that reported a file it never wrote. The minutes are cheaper.
The closing argument
Mutation testing is an old idea that infrastructure code mostly skips, because no tool knows what a Terraform default or a PowerShell guard promises. Written by hand, one plant per promise, it took a day for fifteen repositories and found four tests that could not fail. The badge is the smaller result. The larger one is that 153 times, something broke on purpose and something noticed. A green run means something again.
Sources
- Inozemtseva and Holmes, Coverage is not strongly correlated with test suite effectiveness, ICSE 2014 (31,000 suites, coverage as a quality target), accessed September 2026
- HashiCorp, Terraform v1.7 changelog (mock providers in
terraform test, 1.7.0 released 17 January 2024), accessed September 2026 - OpenSSF, Best Practices badge criteria (the test criteria at the passing level), accessed September 2026
- Microsoft, about_CommonParameters (
WhatIfand its inheritance), accessed September 2026 - OpenSSH, ssh_config manual page (
StrictHostKeyChecking), accessed September 2026 - AWS, Security policies for your Application Load Balancer (the post-quantum recommendation and the defaults by tool), accessed September 2026
- AICPA Trust Services Criteria via ISMS.online, SOC 2 CC8.1, change management, accessed September 2026
- ISO/IEC 27001, 2022 edition, via ISMS.online, Annex A 8.29, security testing in development and acceptance, accessed September 2026
- PIT, Stryker and mutmut, mutation testing for Java, JavaScript and TypeScript, and Python
Discussion
If you have planted a violation in your own tests and watched one survive, or you prove your infrastructure tests some other way, drop a comment below. Counterarguments are welcome and the comment thread is where I respond first. For longer back-and-forth with senior practitioners, join the discussion on Discord.