3334 words
17 min read

Mutation Testing for Infrastructure Code — 153 Planted Bugs, Four Tests That Could Not Fail

By · Solutions Architect · Docker Captain · CNCF Ambassador · IBM Champion
Cyan-lit steel lattice with a red fault inside, mutation testing

Fifteen of my public repositories carried a green CI badge and a README line saying they were verified on every push. Not one of them ran a step that could fail because the code did the wrong thing. Their workflows ran actionlint, tflint, shellcheck, hadolint and a Docker build. Lint is an opinion about text. The badge was being read as a statement about behaviour, and I was the one who had put it there. So I wrote the tests, then ran mutation testing on them by hand: 153 deliberate breaks, one per promise, each required to turn a test red. Four tests stayed green. Two shipped defects surfaced on the way, one of them almost five years old.

This is the same blindness I wrote about in health checks that cannot fail, one layer further out. A health check that cannot go red is a lie told to the scheduler. A test that cannot go red is a lie told to the reviewer, and the reviewer is usually you, six months later, reading a green run as proof.

What a test that cannot fail looks like#

The log cleaner for IIS and Exchange makes three promises. Old logs of the named kinds go, everything else stays, and a WhatIf run touches nothing. The first run caught its WhatIf test in this shape, checking the files and nothing else:

Terminal window
# clean-server-logs.Tests.ps1, the shape the plant caught: the files, not the claim
It 'deletes nothing under -WhatIf' {
& $Script -Path $Root -Days 30 -WhatIf
$OldLog, $OldEtl, $OldBlg, $NewLog | ForEach-Object { Test-Path -LiteralPath $_ | Should -BeTrue }
}

The plant rewrote the script’s own ShouldProcess check so that it always said yes, and the script stopped asking before it deleted. The test stayed green. Remove-Item honours WhatIf by itself, so the files survived either way, and the test was proving a promise PowerShell keeps rather than one the script keeps.

Terminal window
# clean-server-logs.Tests.ps1: the count is the only place a bypassed guard shows
It 'deletes nothing under -WhatIf, and does not claim to have' {
$out = & $Script -Path $Root -Days 30 -WhatIf
$OldLog, $OldEtl, $OldBlg, $NewLog | ForEach-Object { Test-Path -LiteralPath $_ | Should -BeTrue }
$out | Should -Match '^Deleted 0 file\(s\)'
}

With the guard bypassed, the script counts every file it believes it deleted, and the last line goes red. In banking, telecom and hyperscaler support I have watched the same thing happen with every generation of tooling: the instrument is trusted because it was installed, not because it was ever seen to fire.

The difference was one assertion.

Why this keeps happening#

Coverage is the number teams report, and it measures the wrong thing. In what was then the largest study of the question, Inozemtseva and Holmes generated 31,000 test suites for five Java systems of up to 724,000 lines, and used mutation testing to measure how well each suite caught faults. Once the size of the suite was held constant, coverage correlated with that ability only weakly to moderately, and stronger forms of coverage added nothing. Their conclusion was that coverage should not be used as a quality target. A test that runs a line and asserts nothing about it still covers the line.

Infrastructure code gets even less. Mutation testing is old and well understood in application code. PIT does it for Java, Stryker for JavaScript and TypeScript, mutmut for Python. A Terraform module, a PowerShell maintenance script and a CI example meant to be copied have no tool that understands what matters in them, and a random mutation is the wrong instrument anyway. Reversing a comparison somewhere in a module tells you little. Making a database reachable from the internet tells you at once whether the test guarding that promise works.

The badge hides the rest. A workflow that only lints goes green for the same reasons as one that tests, and a reader cannot tell which it was. The OpenSSF Best Practices questionnaire asks eight questions about an automated test suite. Answered honestly for 89 of my repositories on 25 September, the count was 74 with a suite and fifteen without.

A green badge on a repository with no test says lint. It does not say verified.

Risk and blast radius#

Direct exposure is the change record. The Active Directory script in this set disables dormant accounts, and its README tells the reader to run it with WhatIf first. In a regulated environment, that dry-run report is the artefact a change advisory board signs before anything is disabled. SOC 2 CC8.1 asks whether changes are tested before they are implemented, and ISO 27001 control 8.29 asks for security testing in development and acceptance. A green run and a test file satisfy both. Neither asks whether the test could have gone red.

Systemic exposure is the copy. These repositories are templates. A Terraform pipeline or an OS update job is meant to be copied into somebody else’s account and CI, and the copy inherits the tests along with the code. A guard that no test can see ships into every copy as decoration, and the person copying it has no way to know which guards those are. The OS update job is the sharpest case. It logs into a host over SSH and upgrades its packages, so the setting that decides whether it will talk to an impostor matters more than anything else in it, and nothing tested it.

A control nobody has seen fail is a rumour with a badge.

What mutation testing found#

On the first runs five planted violations went unnoticed, and behind them were four tests that could not fail. All 153 are caught now, and so are the six added since. Each finding below is on the ledger with the rule it produced, and all fifteen repositories now pass the OpenSSF badge. The evidence page lists 90 of 90 registered repositories as passing, every answer public and each one naming the file that makes it true.

Terraform: testing a plan with no AWS account#

The eight Terraform pipelines are tested with terraform test, which since Terraform 1.7, released on 17 January 2024, can replace a provider with a mock. Each test plans the configuration with its default variables, against mocked providers, and asserts what would be built. No account, no credentials, nothing created:

# tests/posture.tftest.hcl in the RDS pipeline, shortened: mocked providers, no AWS account
mock_provider "aws" {}
run "the_defaults_are_the_secure_ones" {
command = plan
assert {
condition = aws_db_instance.db_instance_1.storage_encrypted == true
error_message = "db_instance_1 stores data unencrypted"
}
assert {
condition = aws_db_instance.db_instance_1.publicly_accessible == false
error_message = "db_instance_1 is reachable from the internet"
}
}

Between them the eight files make 164 assertions: KMS key rotation, S3 public access blocks, encryption and versioning, the state lock’s encryption and point-in-time recovery, IMDSv2 on instances, invalid-header dropping on load balancers, TLS policy on listeners, flow logs, EKS secrets encryption. On the first run 109 plants broke the promises behind them, and every one was caught.

The assertions were generated from each configuration’s own resources, which is how the more useful output appeared. Generated that way, they could only assert what the defaults deliver, so every place those defaults fall short became a line in the README instead of a silent gap. SSH is open to the whole internet in four EC2 pipelines. Redis is not encrypted at rest or in transit in two. The EKS API endpoint is public. None of these is hidden any more, and none is claimed.

The load balancers show why a list of plants is never finished. Version 1.2.0 moved three of them to AWS’s TLS 1.3 policy from 2021 and called it the policy AWS recommends. It was not. AWS now recommends its hybrid post-quantum policies, and version 1.3.0 moved the three to one of them. The assertion now requires both TLS 1.3 and a post-quantum key exchange in the policy’s name, and two new plants in each pipeline try the old policy and a TLS 1.2 one. Six more plants, 159 in all, and a changelog entry that says the previous release was wrong.

# 17-alb.tf in the GitLab pipeline, inside the HTTPS listener
ssl_policy = "ELBSecurityPolicy-TLS13-1-2-Res-PQ-2025-09"

A test that asserts only the good news is a brochure.

The dry run that reported a file it never wrote#

The Active Directory script disables accounts that have not signed in for a set number of days, stamps them, and moves them to a holding OU. It supports WhatIf properly, through SupportsShouldProcess. Under WhatIf it printed how many dormant accounts it had written to its report.

It wrote no report.

WhatIf sets a preference variable for the whole script scope, and every cmdlet that supports ShouldProcess inherits it, including New-Item and Export-Csv. The account changes were correctly skipped. So was the report, the one thing a dry run exists to produce. The line after it announced a file that did not exist. A two-parameter fix exempts those two cmdlets from the dry run:

Terminal window
# disable-inactive-users-active-directory.ps1: the report is written even under -WhatIf
New-Item -Path $LogFolder -ItemType Directory -Force -WhatIf:$false | Out-Null
$report | Export-Csv -LiteralPath $logFile -NoTypeInformation -Encoding UTF8 -WhatIf:$false

The Pester suite that found it runs the script against a stand-in ActiveDirectory module that records every call instead of making it. That is what makes an AD script testable on a CI runner with no domain. Its tests assert the sequence of changes the script would make, in order, and that WhatIf makes none of them. For three weeks, since the script learned WhatIf on 3 September, the dry run had been a sentence with no file behind it.

The other defect was older. In the terminal tree, the Ctrl-C handler called tput reset, which wipes the scrollback and, on a terminal that does not answer it, stalls long enough for the script to be killed before the cursor comes back. It had been there since the first commit, in December 2021, and a headless test that runs the tree under util-linux script found it on its first run. The handler now restores the colours and the cursor, clears the screen, and exits with the signal’s code:

Terminal window
# terminal-christmas-tree.sh: give the terminal back on Ctrl-C
restore() { tput sgr0; tput cnorm; clear; }
trap 'restore; exit 130' INT
trap 'restore; exit 143' TERM

Both had shipped. Both surfaced the first time a test asked what the script does, not what it prints.

Four tests that could not fail#

The plants did their work on the tests, not the code. Five of them missed on the first runs, from four tests, and each of those tests looked right and proved nothing.

TestWhat the plant brokeWhy the test stayed greenWhat it asserts now
Log cleaner, -WhatIfThe script’s own ShouldProcess guardThe delete cmdlet honours the dry run on its ownThe count the script reports
Mattermost retentionA code path the test never reachedThe test never ran that pathA 20-day-old post that a 30-day retention keeps
Terminal tree sizeThe size of the treeThe check read the drawn output as a wholeThe widest row drawn between cursor moves
OS update over SSHThe empty-scan guard before the loginStrict host key checking refused the host anywayAn impostor takes the target’s name, and the job must refuse

The first is the WhatIf test above. In the Mattermost retention test, one plant broke a code path the test never reached. It was replaced by a plant the test could see, and by a 20-day-old post that a 30-day retention has to keep, because a retention job that deletes too much is the expensive failure, not the one that deletes too little. The terminal tree’s size check read the drawn output as a whole, and a plant that shrank the tree passed it. It now measures the widest row drawn between cursor moves.

The fourth is worth the whole exercise.

The host key check that sat behind a second lock#

Two example jobs ship with the OS update pipeline, one for GitHub Actions and one for GitLab CI, and both SSH into a host and run apt-get upgrade. Each scans the host key, refuses to continue when the scan comes back empty, and then connects with strict host key checking on. The first versions, years ago, connected with it off.

Terminal window
# .github/workflows/00-os-update.yml.example, shortened: the scan, the guard, the lock
ssh-keyscan -H "$EC2_HOST" > ~/.ssh/known_hosts 2>/dev/null
test -s ~/.ssh/known_hosts || { echo "::error::no host key from $EC2_HOST"; exit 1; }
ssh -i ~/.ssh/id_ed25519 -o StrictHostKeyChecking=yes -o BatchMode=yes \
"$SSH_USER@$EC2_HOST" 'sudo apt-get update && sudo DEBIAN_FRONTEND=noninteractive apt-get upgrade -y'

Two plants removed the empty-scan guard, one in each example. The test stayed green.

With the guard gone, strict host key checking refuses the unknown host anyway, so the job still stopped and nothing a test could see had changed. The guard was a second lock behind the first. Testing it proved nothing, because the thing that mattered was the first lock, and nothing tested that.

Strict host key checking exists for one case. The host that answers is not the host whose key you scanned. So the test now builds that case. The runner scans the real target’s key, then the target’s network name moves to an impostor container with freshly generated host keys and the same account, and only then does the job’s SSH step run:

Terminal window
# The target's name moves to the impostor between the key scan and the login.
swap_to_impostor() {
docker network disconnect "$RUN" "$RUN-target" >/dev/null
docker network connect --alias target "$RUN" "$RUN-impostor" >/dev/null
}

The job has to fail, and nothing may run on the impostor. Now the plant that matters turns strict host key checking off. With that plant in place the job logs straight into the impostor and upgrades it. The test goes red, and the run reports the plant as caught. That is a man-in-the-middle test running on a free GitHub runner inside that job, for a pipeline that is copied into other people’s CI.

A security control is tested when an attacker’s move is shown to fail against it. Removing the control’s neighbour is not that.

Framework: promise, plant, and a run that has to go red#

The fleet’s own engine has worked this way since 7 September, when a health check that could never fail made it a rule that every check ships with the test that plants its violation. There are about four hundred of them, now public. The fifteen repositories were where the rule had not reached. Any repository can take the same discipline in three steps.

Layer 1: write the promises and their plants (week 1)#

Write down what the repository promises, one sentence per promise, in the terms a reader would use. “The database is not reachable from the internet”, not the name of the setting that makes it so. For each promise, write the smallest edit that breaks it. That edit is a plant. In a large organisation, start with the promises whose failure would be an audit finding rather than a bug: who can reach the data, what gets deleted, what a dry run records, and which host a job will log into. Every tested repository here keeps them in one tab-separated file beside its tests. Each line names a file, a piece of text in it, what to put there instead, and a sentence saying what the test must notice. This line from the RDS pipeline, shortened here to the lines that change, flips the default for public accessibility:

# tests/plants.tsv: file, old text, new text, what the tests must notice
"00-variables.tf" "type = bool\n default = false" "type = bool\n default = true" "publicly_accessible flipped in the default of rds_publicly_accessible"

The fields are JSON strings separated by tabs, so a plant can span lines and carry quotes without an escaping scheme of its own.

Owner: the repository’s maintainer.

Layer 2: break each promise on its own copy (week 2)#

The runner that reads the plants is about eighty lines of Python. Script repositories share one, and the Terraform pipelines use a variant that runs terraform test in a pinned container. Its core runs the suite on an untouched copy, then applies each plant to a fresh copy and counts what stayed green:

# tests/plant-violations.py: the untouched copy first, then one plant per copy
work = copy()
r = run(cmd, work)
if r.returncode != 0:
sys.exit("the untouched code does not pass its own tests")
missed = stale = 0
for fname, old, new, what in plants:
work = copy()
path = os.path.join(work, fname)
text = open(path, encoding="utf-8").read()
if old not in text:
print(" STALE %s: the text this plant breaks is no longer in %s" % (what, fname))
stale += 1
else:
open(path, "w", encoding="utf-8").write(text.replace(old, new, 1))
if run(cmd, work).returncode == 0:
print(" MISSED %s" % what)
missed += 1

The untouched copy has to pass. Otherwise every plant “fails” for a reason that has nothing to do with the plant, and the run reports a perfect score over a broken suite. And every plant runs on its own copy, because two violations at once can mask each other.

Owner: platform engineer.

Layer 3: run the plants on every push, and fail on a stale one (week 3 onward)#

A plant whose text is gone is stale, and stale fails the run. Code changes. A plant that silently stops applying is a test that silently stops being tested, which is the exact failure this whole exercise exists to catch. So the plants run in CI on every push, beside the tests:

# .github/workflows/verification.yml in the log cleaner
# A test that cannot fail is not a test: each promise is broken on a
# copy of the script, and the run fails if the tests stay green.
- name: The tests, shown each violation they exist for
run: python3 tests/plant-violations.py -- pwsh -NoProfile -NonInteractive -File tests/run.ps1

When a plant survives, fix the test first. The survivor is information, not noise.

Owner: platform engineer, with security owning the plants that guard access: public endpoints, TLS policy and SSH host keys.

Tradeoffs#

The plant runs are not free. Each plant rebuilds or replans on a fresh copy, so a repository with eighteen plants runs its suite nineteen times. On GitHub-hosted runners the slowest here, the OS update pipeline with two container images, spends about three minutes per push on its test and plants together. Public repositories pay nothing for Actions minutes. A private monorepo would want the plants on a schedule or on changes to the tested files, not on every push.

Skipping them costs what this post already showed: a dry run that could have been signed for three weeks with no report behind it, and an SSH guard whose only test proved nothing.

Three minutes a push, against a dry run that reported a file it never wrote. The minutes are cheaper.

The closing argument#

Mutation testing is an old idea that infrastructure code mostly skips, because no tool knows what a Terraform default or a PowerShell guard promises. Written by hand, one plant per promise, it took a day for fifteen repositories and found four tests that could not fail. The badge is the smaller result. The larger one is that 153 times, something broke on purpose and something noticed. A green run means something again.

Sources#

Discussion#

If you have planted a violation in your own tests and watched one survive, or you prove your infrastructure tests some other way, drop a comment below. Counterarguments are welcome and the comment thread is where I respond first. For longer back-and-forth with senior practitioners, join the discussion on Discord.


Vladimir Mikhalev

Docker Captain  ·  IBM Champion  ·  AWS Community Builder

The Verdict — production-tested analysis on YouTube.

The Verdict

Inconvenient truths about shipping in the AI era

Container security, platform engineering, and the agentic shift — tested in production, argued without the hype. The verdict reaches your inbox the moment there's one worth sending.

Related Posts

Same category
  1. 1
    Docker Health Checks That Cannot Fail — Seventeen Hours Green
    DevOps & Cloud · A Docker health check whose pattern matches the shell running it can never go red. How to find the checks in your fleet that never failed, and prove each one.
  2. 2
    Terraform MCP server GA: the Apply Gate your auditor will ask about
    DevOps & Cloud · HashiCorp's Terraform MCP server is GA and IBM Bob can write production IaC. ENABLE_TF_OPERATIONS separates a safe assistant from autonomous apply.
  3. 3
    Docker supply chain hardening — from Scout D to OpenSSF 7.8 on a 730K-pull image
    DevOps & Cloud · How I hardened a 1M-pull public Docker image from Scout grade D to OpenSSF Scorecard 7.8 — multi-stage build, cosign, SLSA provenance, non-root default.
  4. 4
    Cloudflare Web Analytics on Astro — Why Removing GA4 Unlocked Lighthouse 100
    DevOps & Cloud · How removing Google Analytics 4 from an Astro site unlocked Lighthouse 100, why Cloudflare Web Analytics replaced it, and what the tradeoffs actually cost.

Random Posts

Random
  1. 1
    Basic Setup of Windows Server 2012 R2
    SysAdmin & IT Pro · Step-by-step guide to setting up Windows Server 2012 R2 - hostname, RDP access, time zone, static IP, domain join, and system locale configuration.
  2. 2
    Choosing Between Docker Swarm and Kubernetes for Container Management
    DevOps & Cloud · Compare Docker Swarm vs. Kubernetes for container orchestration. Explore key differences in scalability, security, networking, and DevOps integration.
  3. 3
    Install OTRS Using Docker Compose
    Self-Hosting · Learn how to deploy OTRS Helpdesk with Docker Compose, secured by Traefik and Let's Encrypt. Step-by-step guide for Ubuntu-based self-hosted ticketing.
  4. 4
    Mastering Terraform Contains and Strcontains Functions
    DevOps & Cloud · Learn how to use Terraform's contains and strcontains functions for better logic control in IaC. Includes practical DevOps examples and best practices.
Mutation Testing for Infrastructure Code — 153 Planted Bugs, Four Tests That Could Not Fail
https://heyvaldemar.com/mutation-testing-infrastructure-code-planted-violations/
Author
Vladimir Mikhalev
Published
2026-09-27
License
CC BY-NC-SA 4.0