hraness
Theme
Appearance

planted bugs: how to test the tests

plant a bug and see whether the tests fail

by hraness · drafted with ai assistance

the rest of this lesson is free: add your email to keep reading.

To learn whether a test suite would catch a bug, plant one and see whether the suite fails. That is mutation testing. A passing suite tells you nothing about the bugs nobody wrote yet, and line coverage measures which lines ran, not which mistakes would be caught.

a green suite can guard nothing

Every check in a pipeline is itself code: the assertion, the invariant, the policy rule. If the check is broken, through a wrong comparison, an empty condition, or a variable nobody reads, the suite reports green while guarding nothing. The bug then slips past the tests quietly and shows up in production.

Mutation testing asks whether the tests fail on deliberately wrong code. A mutant is a seeded defect: flip the comparison, drop the guard, swap the order. If the suite still passes, the mutant has “survived,” and that survivor marks a measured hole in the tests.

three kinds of planted bug

Useful mutants look like real mistakes rather than random syntax changes. The portfolio uses three kinds.

Mutant specs and models. vhalla runs its TLA+ specifications against deliberately wrong protocol variants to show the spec can fail at all. A spec that never rejects a mutant asserts a property too weak to catch the bugs it was written for.

Boundary mutants. Gobstopper seeds defects into its admission and boundary checks. The suite runs against a version of the check with the condition inverted or the limit removed, and it must fail.

Field mutations. ALGAL’s contract tests mutate manifest fields: they swap a bound, drop a required field, or widen a capability, then confirm ALGAL rejects exactly those changes. The mutants pin down what ALGAL accepts, beyond the happy path.

mutants inside the test suite

The portfolio keeps mutants in the test suite instead of in a separate tool that runs now and then. A mutant harness writes the defective variant (a patched config, a flipped flag, a wrong constant), runs the suite or the specific check against it, and asserts that it fails. The test then shows that the guard does real work.

Each surviving mutant names a place where wrong code would pass. Each killed mutant shows the check catches that kind of mistake, so coverage comes to mean mistakes caught rather than lines executed.

what a killed mutant does not prove

Mutants measure the tests. Killing every mutant shows the suite catches the kinds of bug you planted, and bugs of other kinds stay invisible, so mutation testing adds to the other techniques in this series without replacing them.

Mutants exist only where someone thought to plant them. A set that covers easy flips (off-by-one, inverted booleans) and misses harder ones (wrong order of operations, a missing fsync, a dropped limit) measures the shallow holes and leaves the deep ones unmeasured. Plant mutants where the system actually fails, and name the set in the test.

A surviving mutant needs judgment before it becomes a finding. Some are equivalent mutants, where the change does not alter behavior, and dismissing them takes a person or an agent who understands the code. A mutant suite that runs longer than the tests themselves tends to get skipped, so keep it small and aimed at the checks that decide whether a bug ships.

keep reading: free for subscribers

the rest of this lesson is free. enter your email to subscribe, and every subscriber lesson unlocks in this browser.

already subscribed? enter the same email to unlock.