A property test states a law and then tries to break it a few thousand times. Where an example test says “for this input, expect this output,” a property test says “for every input of this shape, this relationship holds”: parse(serialize(x)) == x, sorted output is a permutation of the input, decode never panics. For code that reads foreign input, that difference decides whether the bug is found in CI or in production.
why laws find parser bugs
A parser’s bug rarely sits at an input a developer thought to write down. It sits at the input that is almost valid: the IDN that normalizes oddly, the path with a drive-letter segment, the JSON with a duplicate key. A generator produces those inputs systematically. A hand-written example list almost never does.
Across the portfolio, the convention pairs concrete regression tests with property tests at four boundaries:
- Parsing. Round trips hold, and adversarial input is rejected.
- Resolution. Paths and identities stay inside their scope.
- Ordering. Output order depends on the input, never on map iteration.
- Confinement. Nothing escapes its directory or scope.
how the repositories use them
GhostGet’s URL checks are the clearest case. Candidate URLs run through a comparedParse that compares the TypeScript acceptance predicate with the Rust url crate, which serves as an oracle. Two independent implementations of one standard that disagree expose a failure neither can see in itself. The property is “the two agree,” and every known deviation is written down with its reason.
Oh’s parity suite applies the same idea to its memory kernel. The kernel’s operations have a documented contract, and the property tests generate operation sequences that must behave identically against the model. Accounts and design-kit test with size-limited generated inputs for a related reason: their catalog types reject values they cannot name, and the tests check that rejection across generated inputs instead of a few chosen ones.
xcb’s task history and aicharts’ benchmark projections use ordering and determinism properties. The same input history must project to the same output, byte for byte, so replaying a run also checks that it is deterministic.
keeping every counterexample
A random generator that finds a bug once will probably never find it again, which leaves the fix unverifiable. The repositories record the case instead. GhostGet keeps seeds/corpus.json, where each entry names a seed and a path, so the counterexample replays deterministically inside the suite.
That turns a flaky failure into a permanent test. CI once found a disagreement between the runtime’s URL check and the Rust oracle on an input with a backslash-normalized drive-letter segment. The fix corrected the exclusion rule in comparedParse rather than patching the one case, and the seed is now pinned so the input replays on every run.