In short: A single holdout window measures one quarter's demand environment rather than a model's skill, and the ranking flips when you move the cut-off. Four design choices decide what the harness actually measures: where the origins sit and how far apart, whether the window expands or slides, whether parameters are refit at each origin the way production refits them, and how per-item errors are pooled into one number. Overlapping forecast windows are correlated, so forty weekly origins on a thirteen week horizon carry closer to three independent reads. Budget your comparisons before you run them, because the best of forty configurations on one holdout is expected to look good by luck.
The challenger won by four percent on a twelve week holdout ending in March. Someone re-ran the same comparison in July against a holdout ending in June, and the champion won by two. Neither run had a bug, both numbers were computed correctly, and the team spent a fortnight arguing about which one to believe.
They were both measurements of something real. The something was mostly the demand environment in one particular quarter. A backtest is a measuring instrument, and the design of the instrument decides what it measures. Most of those design decisions get made once, by whoever wrote the first evaluation script, and are never looked at again.
The three numbers that define an evaluation
Every out-of-sample test has an origin, a horizon and a gap, and the third one is usually wrong.
The origin is the last date the model is allowed to see. The horizon is how far ahead it forecasts, and errors should be reported per step rather than averaged across the whole horizon, since a model that is excellent at one week ahead and poor at thirteen is a different proposition from one that is mediocre throughout. How accuracy decays with horizon, and what to do about it, belongs to BB15.
The gap is the number of periods between the origin and the first forecast period, and it exists because data arrives late. If point of sale files land with a nine day lag, and your planning cycle locks on a Wednesday, then the model running in production on any given Wednesday has visibility of demand up to roughly a week and a half earlier. A backtest with a zero gap gives the model information the production model will never have, and the difference shows up as accuracy that evaporates on go-live. Set the gap to the real latency of the slowest input, per input, and let features carry their own availability lag rather than assuming everything arrives together. The related question of features that encode information from the future is covered in J29.
Tashman's 2000 review in the International Journal of Forecasting is still the clearest statement of the vocabulary here, and it separates fixed origin evaluation, where you cut once and forecast forward, from rolling origin evaluation, where the cut moves through the data and you accumulate errors from many starting points. The fixed origin version is what almost everyone builds first, and it is what produced the March and June disagreement above.
How the window moves and when you refit
Once the origin moves, two further choices appear, and both of them change the answer.
Expanding or sliding. An expanding window trains on everything up to each origin, so the training set grows as the evaluation walks forward. A sliding window keeps a fixed length, dropping the oldest period as it adds a new one. Expanding matches a production system that refits on full history. Sliding matches one that caps training at three years, and it has the secondary property of making early and late origins comparable, since every fit sees the same amount of data. When your evaluation shows a model improving over successive origins on an expanding window, you cannot tell whether the model got better or the training set got bigger.
Refit or roll forward. Tashman distinguishes updated from non-updated forecasts, meaning whether model parameters are re-estimated at each new origin or fitted once and carried forward with only the state updated. This is the choice most often made on the basis of compute cost, and it should be made on the basis of what production does. A gradient boosted model with four hundred features, refit at every weekly origin in the backtest but refit quarterly in production, has been evaluated as a product you are not going to ship. The error runs in the flattering direction, and on fast-moving assortments it is large.
The rule is easy to state and unpopular to implement. The backtest should replicate the production cadence: the same refit frequency, the same training window policy, the same feature availability at each origin. Anything the harness does that production will not do is a subsidy to the model under test.
How many origins before a difference means anything
Here is the arithmetic that usually surprises people. Three years of weekly history gives 156 observations. Reserve 104 weeks for the first training set and forecast a thirteen week horizon, and you get 40 weekly origins, which is a comfortable-sounding number.
The overlap between those origins takes most of it away. Origin t and origin t+1 share twelve of their thirteen forecast periods, so their errors are almost the same errors. Independent blocks are spaced roughly a horizon apart, which leaves something closer to three non-overlapping reads of the model's behaviour. Three reads is enough to notice that a model is broken and nowhere near enough to certify a small improvement.
Two things buy back statistical power. The first is the cross-sectional dimension: five hundred items evaluated at the same origins gives many more error observations, though item errors are correlated through common shocks, so the effective sample is smaller than the raw count suggests. The second is looking at the sign of the difference rather than its size. Count the fraction of items where the challenger beat the champion at each origin, and check whether that fraction stays on the same side of a half across origins. A model that wins on 58 percent of items at every one of your origins is more convincing than one that wins by a large margin at a single origin and loses at the next.
If you want a formal test, Diebold and Mariano's 1995 procedure in the Journal of Business and Economic Statistics compares two forecasts through the series of loss differences, and Harvey, Leybourne and Newbold published a small sample correction in the International Journal of Forecasting in 1997 that matters at the sample sizes planning teams actually have. With overlapping h-step forecasts, the loss differential series is autocorrelated, so the variance has to be estimated with a lag window of at least h minus one. Skipping that correction is the standard way to manufacture significance.
Pooling errors across items decides the winner more often than the models do
A backtest produces an error for every item, at every origin, at every step. Collapsing that cube into one number is a modelling decision dressed as a reporting decision.
An unweighted mean across items lets the long tail decide. Percentage errors on items selling two units a week have tiny denominators and enormous values, and a handful of them will dominate the average regardless of what the models did on the rest of the catalogue. Volume weighting swings the other way, and on a typical assortment the top twenty items will settle the comparison between them. Both pooling schemes are defensible and they routinely disagree, which is why the answer to "which model is better" is genuinely ambiguous until someone says what the model is for.
Match the pooling to the decision. If the backtest is choosing a default for thousands of unattended tail items, weight by item count and accept that the big movers are irrelevant to that question. If it is choosing the model for the items a planner touches every week, weight by volume or contribution margin and stop pretending the tail matters. Running both, and reporting them separately, is better than arguing about which single number is correct. Which error metric goes inside the pooling, and why a scaled measure behaves differently from a percentage one, is BB2's territory.
Report a distribution rather than a point. The three figures worth putting on the same page are the median per-item error, the win rate against the incumbent, and the ninetieth percentile of error. A challenger that improves the mean while losing on sixty percent of items has found a few series it is very good at, which is useful information and a bad reason to replace a default.
The comparison budget
Every configuration you try on the same evaluation data is a draw from a distribution, and the maximum of many draws is high by construction. Bailey, Borwein, Lopez de Prado and Zhu made this point sharply in the Notices of the American Mathematical Society in 2014, in a finance setting: with enough trials, an impressive backtest carries almost no information about future performance, because the selection procedure guarantees an impressive result exists.
Planning teams run into this without noticing, because the trials are spread over months. Someone tests four model families, then a few feature sets, then some hyperparameters, then a couple of aggregation levels, all scored on the holdout ending in March. Nobody counted, and the count is in the dozens.
Two habits contain it. Hold back a set of origins that no configuration is scored against until the decision is final, and treat them as the only reportable numbers. Then write the decision rule down before running the comparison, including how large a difference has to be before the champion is replaced. Deciding what counts as a win after seeing the results is how a two percent difference becomes a migration project. The governance around that promotion decision, and what to do when the challenger only wins on part of the catalogue, is T2.
When ordinary k-fold is defensible
Standard k-fold cross validation, where interior blocks are held out and the model trains on data from both sides, is usually treated as forbidden on time series, on the reasoning that training on the future to predict the past cannot mean anything. The prohibition is broader than the evidence supports, and there is a case where the efficiency gain is worth having.
Bergmeir and Benitez showed in Information Sciences in 2012 that cross validation can work on time series predictor evaluation, and Bergmeir, Hyndman and Koo followed it in Computational Statistics and Data Analysis in 2018 with a specific result: for purely autoregressive models whose residuals are uncorrelated, k-fold cross validation is valid and gives a lower variance estimate of error than evaluating on the last block alone. That matters when history is short. An item with sixty weeks of data cannot spare a thirteen week holdout without losing a season.
The conditions are real, though, and they are checkable. Test the residuals for serial correlation before relying on it, and abandon it as soon as the model uses exogenous drivers whose future values you would not have known at forecast time. Promotion calendars, price plans and weather all fail that condition in different ways. Where in-sample residual dependence remains, the interior folds are contaminated by their neighbours and the error estimate is optimistic. Lopez de Prado's purging and embargo scheme from Advances in Financial Machine Learning (2018) is the standard repair, and it amounts to deleting observations either side of each fold boundary so the training and test sets cannot see each other through overlapping windows.
Where this stops
A backtest scores a model against the past regime. If the last two years contained a demand shock, a distribution change or a pricing reset, the harness will reward whichever model handled that particular disturbance, and there is no statistical way to know whether the next disturbance rhymes with it. Rolling origins reduce the problem by averaging over several regimes, and they do not solve it.
The history the harness scores against is also the history your decisions produced. Weeks where you were out of stock recorded the sales you could serve rather than the demand that existed, so a model that predicts constrained sales well can score better than one that predicts demand well. Unconstraining that history is a separate exercise, and D1 covers it.
There is a power ceiling worth being honest about. With three or four independent origins and a few hundred items, a well-designed harness can reliably detect a large difference between models and cannot resolve a one percent difference in scaled error. That ceiling comes from the amount of information in your history rather than from anything wrong with the harness. It also means the search for a one percent gain is not worth funding, since you would not be able to tell whether you had found it, and the accuracy gain that does survive still has to be translated into a decision that changes before it is worth anything, which is what D7 measures.
Take last quarter's model comparison, move the origin back four weeks, run it again, then move it back eight and run it a third time. If the winner changes, you have learned more from twenty minutes of compute than the original bake-off told you.