In short: Planning system testing breaks when it is written in transactional style, because a planning output is a computed recommendation rather than a value anyone can look up. Three layers cover it: arithmetic identities that must hold exactly, properties that must be true whatever the numbers are, and comparative tests against a known baseline. Testing on clean data is the usual way to arrive at go-live with a system that has never met the conditions it will operate under. Agree the exit criteria for a parallel run before it starts, because afterwards every argument becomes one about whether a given difference matters.
User acceptance testing passed. Three hundred and forty test cases, all green, signed off by the business. Two weeks after go-live a planner notices that the forecast for a large customer looks about twenty percent low, and by the time anyone works out why, four weeks of replenishment has been built on it.
Nothing was wrong with the testing effort. The problem is that planning systems resist the test method most people bring to them. A transactional test has an expected result: post this invoice, expect this ledger entry. A planning test usually does not, because the output is a computed recommendation whose correctness is a matter of judgement rather than a value someone can look up. Writing test cases in the transactional style produces cases that confirm the system ran, which is not the same as confirming it was right.
The way through is to stop looking for a single expected value and test in three separate layers, each of which does have a definite answer.
Layer one: identities that must hold exactly
A large amount of what a planning system does has arithmetic that is exactly checkable, and this is where the cheapest and most reliable tests live.
Aggregation adds up. The sum of the item-level forecasts equals the product family forecast, for every family, for every period. If reconciliation is applied, the identity holds after reconciliation rather than before.
Disaggregation sums back. Enter a number at an aggregate level, spread it down, and check the children sum to the parent. Include the case where the disaggregation basis is zero for some children, which is where most systems behave in a way somebody has to decide about.
Unit conversions are reversible. Convert to base and back and get the original figure, to a stated rounding.
A copy is a copy. Copy a version or a scenario and compare it cell for cell against the source.
Netting is conserved. On-hand plus scheduled receipts minus requirements equals the projected balance, period by period, with no leakage.
These tests are boring, they are automatable, and they catch a class of defect that is otherwise found by a planner six weeks later. Write them as queries against the output rather than as manual steps in a test script, and run them on every build.
Layer two: properties that must be true regardless of the numbers
The second layer tests relationships rather than values, which is what lets you test a calculation whose correct answer you do not know.
Raising a service level must not lower the safety stock. Adding demand must not reduce a requirement. Extending a lead time must not shorten a planned order's release date. Removing a constraint must not worsen the objective in a constrained optimisation. Each of these is a property that should hold for every item in every dataset, and each is checkable by running the calculation twice with one input varied.
The most valuable property in this layer is determinism. Run a plan twice against frozen inputs and compare the outputs cell by cell. The correct answer is zero differences.
When it is not zero, the cause is almost always one of four things: a transformation reading the system date, an unseeded random number in a heuristic or a sampling step, a parallel aggregation whose floating-point order varies between runs, or an input that was not as frozen as you thought. All four are findable in an afternoon once you know the count is non-zero, and all four are close to unfindable if you do not.
Put a number on why this matters. If a determinism defect changes 0.4 percent of cells by a small amount on each run, and your model has 80 million cells, that is 320,000 cells moving every time the plan runs. Most will be immaterial. Some will cross a reorder point, and the resulting orders will appear and disappear between runs, which is exactly the behaviour that makes planners stop trusting a system.
Layer three: comparative tests against a known baseline
The third layer is where the judgement lives, and the trick is to convert judgement into comparison.
Build golden datasets: a fixed input set with stored outputs that were reviewed and accepted at a point in time. Every subsequent build runs against them and reports differences. A difference is not automatically a failure, since a deliberate improvement will produce differences too. It is a prompt for someone to look and decide.
Report differences by magnitude rather than by count, because a count treats a rounding difference and a doubled forecast identically. A useful report has three lines: the number of cells differing at all, the number differing by more than the tolerance, and the largest absolute and relative differences with the item identified.
Choose the tolerance per measure. A forecast quantity might tolerate 0.5 percent. A safety stock might tolerate 1 percent. A planned order date should tolerate nothing, because a date is discrete and a one-day shift is a real change.
Keep the golden datasets small enough to run on every build and varied enough to be worth running. Three or four of them covering different conditions works better than one large one: a fast-moving seasonal set, an intermittent set, a set with a promotion in the window, and a set containing the master data defects. When a change breaks only the intermittent set, you know where to look before opening anything.
Test data that resembles production
Testing on clean data is the most common way to arrive at go-live with a system that has never met the conditions it will operate under.
The requirement is a subset that keeps the shape of production rather than a random sample. A 5 percent random sample of items loses almost everything interesting: the long tail is under-represented, the items with unusual master data are probably absent, and the hierarchy is full of holes.
Sample by segment instead. Take every A item. Take 20 percent of B items. Take 2 percent of C items, chosen to include the intermittent ones. Take every item that appears on your master data defect list. Take every item involved in a known edge case: a unit of measure change, a mid-year hierarchy move, a discontinued predecessor, a co-packed product. Then take every location and every customer that any of those items touch, so the hierarchy is complete for the items in scope.
Work out the size before you build it. Suppose production has 800 A items, 6,000 B and 45,000 C. The rule above gives 800 plus 1,200 plus 900, which is 2,900 items, plus a few hundred deliberately awkward ones. That is a dataset small enough to run quickly and large enough to contain the problems.
Mask what needs masking, and mask consistently, so that the same customer maps to the same pseudonym in every table. Inconsistent masking breaks joins and produces test failures that have nothing to do with the system.
The part people leave out is the ugliness. Production master data contains items with no lead time, locations with no calendar, and units of measure that convert wrongly. If the test dataset has been cleaned, the system will meet those conditions for the first time in production. Keep a deliberate sample of bad records and assert what the system should do with each.
The parallel run
A parallel run is where the old and new systems operate on the same inputs for a period and the outputs are compared. It is expensive, it is worth doing, and it goes wrong when the exit criteria are agreed after it starts.
Settle five things in writing beforehand.
Duration, expressed in cycles. Six weeks means nothing on a monthly process. Three complete cycles is a reasonable minimum, because the first is contaminated by learning and the second by fixes.
What runs in both. Usually the forecast and the supply plan, sometimes only the forecast. Anything not in scope should be stated, so that a difference in something out of scope does not become an issue.
Who arbitrates a difference. A named person with the authority to declare a difference explained and closed. Without this the list only grows.
Numeric exit criteria. Something of the form: fewer than 1 percent of item-weeks differ by more than tolerance, no unexplained difference exceeds a stated value, and every root cause identified has been either fixed or accepted in writing.
What happens when the old system is wrong. It will be. Agree in advance that a difference traced to a defect in the old system closes as resolved rather than requiring the new one to reproduce it.
Expect the difference count to fall in steps rather than smoothly, because differences cluster into a small number of root causes. A first run producing 1,400 differences above tolerance out of 30,000 item-weeks, which is 4.7 percent, will typically resolve into five or six causes, and fixing them takes the count under 1 percent in one or two cycles. A count that stays flat across cycles means the causes are not being found, and that is the signal to stop and change approach rather than to extend the run.
Plan the workload honestly. A parallel run means somebody produces two plans every cycle and reconciles them, which is roughly double the work for the duration. If that resource has not been allocated, the parallel run degrades into the new system running unattended while everyone uses the old one.
Decide up front which system's output the business actually executes on during the run, and say it out loud. Running both and executing on neither consistently is the worst outcome, because a difference then has no consequence and nobody investigates it with any urgency. The usual arrangement is that the old system remains authoritative until the exit criteria are met, with a stated date on which that reverses.
The limits
None of these layers tells you whether the plan is any good. They tell you that the system computes what it was configured to compute, consistently and reproducibly. Whether the configured method produces better decisions than what you had is a different question, answered by scoring against outcomes on held-back data, and E4 covers how to design that so the result means something.
The golden dataset approach also decays. Datasets age, the business changes, and eventually the stored outputs describe a model nobody uses. Refresh them on a schedule, and treat a golden dataset that nobody has updated in a year as untested rather than as passing.
And there is a limit on determinism worth knowing about. Some legitimate methods are stochastic: a simulation, a metaheuristic, a sampling-based optimiser. For those, the correct test is that the result is stable within a stated tolerance across seeds, and that the seed is recorded with the output so a specific run can be reproduced. Insisting on exact determinism where the method is genuinely stochastic will send somebody looking for a defect that does not exist.
Start by running your plan twice on frozen inputs and counting the cells that differ. That single number takes an hour to produce, and if it is not zero you have found something worth knowing before any of the harder testing begins.