Why golden-master tests come before the first agent edit

View as Markdown
Ask AI
Share

Legacy systems rarely have a specification. The code is the specification, including quirks that other systems now depend on. A batch job that rounds a fee slightly differently from the documented rule has been doing so for twenty years, and every downstream reconciliation has quietly adjusted to it. That is why our approach starts by recording what the current system does, not by rewriting it.

Pin behaviour, then change it

Golden-master tests, also called characterization tests after Michael Feathers’ Working Effectively with Legacy Code, capture the exact outputs of the existing system for a wide set of inputs. They do not judge whether the behaviour is right; they record what it is. An agent’s migrated module must reproduce those outputs before a pull request is even opened.

This gives reviewers a clear contract. They can focus on design and readability instead of wondering whether a rounding rule from decades ago still behaves the same way.

Why agents need this more than people do

A human developer migrating a module carries a lot of implicit caution. They notice when something looks odd, ask a colleague, and test the edge they are worried about. An agent is fast and persistent, but its sense of “done” comes from the signals it is given. If the only signal is “it compiles”, it will stop at “it compiles”.

Golden-master tests turn “behaves exactly like the old system” into a signal the agent can check on every iteration. When a test fails, the agent sees a concrete difference: this input, this expected output, this actual output. That is something it can reason about and fix. Without the tests, the same mistake would surface weeks later in a reconciliation report.

How we plan to capture the master

The quality of a golden master depends on the inputs. We see three useful sources:

  • Recorded production runs. Input files and output files from real batch runs, with personal data masked or synthesized. These cover what actually happens.
  • Branch-guided inputs. The Atlas tells us which conditional paths exist in a program. We generate inputs to reach branches that recorded runs never touch, such as year-end processing or rare error codes.
  • Boundary values. Maximum field lengths, zero and negative amounts, dates around month ends and leap years. Legacy numeric handling is where most silent differences hide.

Exact comparison, with explicit exceptions

By default, outputs are compared byte for byte. Some differences are expected, for example timestamps, generated identifiers or a sort order that was never guaranteed. Those exceptions are written down as named, reviewed rules rather than being hidden inside a loose comparison. A reviewer can see exactly which differences were allowed and why.

What golden masters do not do

A golden master freezes behaviour, bugs included. That is intentional. Migration and bug fixing are separate changes, and mixing them makes both harder to review. When a team finds a bug during migration, we want it recorded, migrated faithfully, and then fixed in a separate, clearly labelled pull request.

Golden masters also do not cover everything. Interactions with external systems, timing-dependent behaviour and concurrency need their own treatment. We will write about those as we work through them with design partners.