Write the system twice, and each implementation can check the other. That is the cheapest answer to a hard testing problem: how do you test a whole runtime (a parser, an evaluator, a replay engine) when the only source of expected output is the system itself? Two implementations of the same contract, run on the same inputs, must agree byte for byte. When they do not, the disagreement points at a bug or at a gap in the contract.
tests need an oracle
An oracle is whatever tells a test what “correct” means. For most code, the oracle is the programmer’s intent written as examples. A deterministic system, such as a manifest interpreter, a state machine, or a projection, allows a stronger oracle: any second faithful implementation. Run the same input through both and require bit-for-bit agreement.
Independent implementations tend to make different mistakes. They share a mistake mainly where the contract is ambiguous, and in that case the contract is what gets fixed.
parity in ALGAL and Oh
In the portfolio, parity means driving the same artifacts through both runtimes and getting identical run receipts, the records each run writes of its inputs and outputs. ALGAL, a language and runtime for agent programs, is the fullest case. Its TypeScript reference runtime and its Rust kernel each execute the bundled example organisms (ALGAL’s name for its programs), and the parity harness compares their receipts byte for byte. Verification runs in both directions too: a receipt produced by the TypeScript CLI must replay and verify under the Rust CLI, and the reverse. That tests determinism itself across the language boundary.
Oh’s parity suite applies the same law to its memory kernel, with one difference. Operation sequences run against a model and against the implementation and must agree, so the second “implementation” is a model in the same language.
what a divergence reveals
The mechanics are deliberately dull. A parity suite lists its example set (every bundled organism, every CLI command that emits receipts, every replayable lifecycle), runs each through both implementations, and diffs the normalized output. The suites are split into store, CLI, inference, and application lifecycle parity, so a divergence arrives already located.
Parity catches bugs no single-runtime test can see: float formatting that differs, JSON keys serialized in a different order, a hash computed over slightly different bytes, a boundary condition two implementations read differently from the same spec. Each divergence forces the contract to become more precise, and that precision is much of what the suite produces.