In short: The level of a spot rate is close to a random walk at the horizons planners buy at, so a flat carry-forward of today's spot is the benchmark any model has to beat under rolling-origin validation. The forecastable half sits on the supply side, where the newbuild orderbook publishes capacity two to three years ahead and Kalouptsidi's time-to-build work explains why that capacity keeps arriving after the rates that ordered it have gone. Your invoice and the published index diverge by a residual you can measure in an afternoon, and that residual decides whether index-linking protects you. What a planner can act on is a calibrated interval, because every decision it feeds turns on the width.
It is late September and finance wants one number. Freight cost per unit for next year, to two decimal places, to sit inside standard cost. You have twelve months of paid invoices, a spot index that has moved by half since spring, and a carrier account manager who says the market firms after Chinese New Year. Whatever you write down becomes the baseline every freight variance for the next four quarters gets measured against.
UNCTAD's Review of Maritime Transport in 2021 estimated that the container rate surge then under way would lift global import price levels by around eleven per cent. At that scale a freight assumption becomes a pricing input, which is a separate decision (C2).
The argument that follows a request like this is usually about the number. A better use of the week is to separate the parts you can genuinely forecast at the horizon you buy at from the parts you are choosing to carry as a range, and to say out loud how wide the range is.
What the forecast has to beat
Start with the benchmark, because most freight rate forecasts are never scored against one. For a series as autocorrelated as a freight rate, the naive model is strong: take today's spot and carry it forward flat. At two to four weeks it is very hard to beat and at three months it is still respectable. Where a settled futures or forward market exists, its curve is the second benchmark, since it aggregates the views of people with money at risk.
Batchelor, Alizadeh and Visvikis tested time series models against exactly this benchmark on Baltic freight rates and forward freight agreements in the International Journal of Forecasting in 2007. The models earned their keep on forward prices, while on spot the random walk was hard to displace at short horizons. Worth re-running on your own lanes before believing a vendor has moved past it.
How you score matters more than which model you pick, and rolling-origin cross validation is the only honest procedure. Step back to each historical decision date, fit on the data that existed then, forecast the horizon you actually buy at, and compare against what happened. An in-sample fit on a rate series looks superb and means nothing, because yesterday's rate explains today's and the model takes credit for knowing it.
Set the bar accordingly. Beating a flat carry-forward by a few per cent of mean absolute error at three months is a real result, and a model claiming to call turning points should show its calls on the last two turns, dated.
Vessel capacity is published two years ahead
Capacity is one of the few genuinely forecastable quantities in this business, because it takes a long time to build. A container ship ordered today arrives in roughly two years, the orderbook and the quarterly delivery schedules are public, and demolition and idling are observable. That gives you a capacity path two to three years out that needs no forecasting, only arithmetic on published data.
Myrto Kalouptsidi built that lag into a structural model of bulk shipping in the American Economic Review in 2014, and the mechanism is the one every shipper has lived through. High rates trigger orders, the ships arrive two years later, and by then the conditions that justified them have turned. Robin Greenwood and Samuel Hanson made the investment side explicit in the Quarterly Journal of Economics in 2015: ship prices and orders move in waves, and high current earnings systematically forecast low subsequent returns, because buyers extrapolate the present. Martin Stopford's Maritime Economics, third edition, 2009, traces the same cycle across a century and a half.
What that gives a planner is a supply path rather than a rate, and turning it into something usable means working with effective capacity, which moves for reasons that have nothing to do with hulls. A routing change sending Asia to North Europe services around the Cape of Good Hope instead of through Suez adds close to a fortnight each way and absorbs ships without carrying a single extra box. Port congestion does the same thing by accident. Slow steaming does it on purpose.
Then compare that path against a demand range and read utilisation rather than price. Rates are a convex function of utilisation: below a threshold carriers compete and rates sit near operating cost, above it the marginal box has no home and the price moves in a way that has no ceiling anyone can name. That convexity is why the distribution of freight rates is right-skewed, and it is the thing to hold in mind when someone asks for a symmetric plus or minus fifteen per cent.
Truckload has the same structure on a shorter lag, where tractor orders and net operating authorities set capacity over months instead of years, and the load-to-truck ratios from DAT Freight and Analytics alongside the Cass Information Systems shipment series are the utilisation read.
The gap between the index and your invoice
Everyone quotes an index. Few teams have measured how well theirs tracks what they pay.
The public indices measure different bundles. Drewry's World Container Index is a composite over a fixed set of major routes. The Freightos Baltic Index, published with the Baltic Exchange, has its own lane definitions and its own rules about which surcharges sit inside the number. The Shanghai Containerized Freight Index, from the Shanghai Shipping Exchange, is built on export cargo out of China and carries that basis with it. Xeneta's XSI is assembled from contracted rates, which makes it a different object again. Two of them can move in opposite directions in the same week without either being wrong.
Your invoice is a stack: base ocean freight, bunker adjustment, peak season surcharge, congestion or equipment imbalance charges when they are running, terminal handling at both ends, and the documentation fees on the lane. An index governs part of that stack, so if the base is sixty per cent of the all-in cost, index-linking the base leaves forty per cent floating on terms the carrier sets.
The measurement takes an afternoon: regress twelve months of paid all-in cost per container on the weekly index for the same lane and week, then read the standard deviation of the residual. That is the basis risk you keep after index-linking, it is usually larger than people expect, and it decides whether an index-linked contract is protection or paperwork.
Why the width matters more than the level
Once the benchmark and the basis are settled, the useful output changes shape. A planner rarely needs to know the rate will be 2,400. They need to know how much of the plan breaks at 3,600, and how likely that is.
Quantile forecasts are the right object, and they need to be calibrated rather than merely wide. Conformalized quantile regression is a reasonable default, because it gives finite sample coverage without assuming the shape of the residual distribution, and freight residuals are visibly non-normal. Rates climb faster than they fall, so the upper tail runs far longer than the lower one, and a method that assumes symmetry prices the wrong side of the risk.
What comes out is a set of numbers finance can work with. Standard cost at the median. A declared freight reserve sized off the eighty-fifth percentile instead of a round ten per cent. A stated rate level at which the mode mix changes. Running one demand plan against three rate paths is a scenario overlay problem, and the overlays need to be versioned copies of one plan rather than three plans that drift apart inside a fortnight.
Then check that the intervals are honest, which is harder than it sounds when annual contracting gives you one real test a year. Score the model far more often than you act on it: every week, on every lane you buy, store the forecast at the horizon you care about and score it later against the outturn. Twenty lanes over two years of weekly origins is a few thousand scored forecasts, correlated enough that the effective sample is much smaller, and still enough to tell whether the model beats a carry-forward. Report pinball loss across the quantiles you publish and a plain calibration count: over the last hundred forecasts, how often did the actual land above your ninetieth percentile? At twenty-five times, the reserve is badly under-sized and no improvement in the median will repair it.
The decisions the forecast is supposed to change
A rate forecast that changes nothing is an expensive newsletter. Four decisions should move with it.
The contract and spot mix. How much volume sits on fixed rate, how much on index-linked, how much stays on spot. The width drives this rather than the median: a wide distribution argues for more fixed cover where the fear is price, and more contracted capacity where the fear is being rolled.
What a fixed contract is really buying. In a rising market the exposure is availability: cargo gets rolled, sailings get blanked, and a paper rate you cannot get equipment against is worth very little. In a falling market, compliance erodes from the shipper's side because spot is cheaper, and minimum quantity commitments get enforced loosely in both directions. Treating a fixed contract as a capacity reservation with a price attached produces better decisions about how much to buy, and running the tender that sets it is a separate exercise (Z1).
When to lock. If the distribution is wide and your volume commitment is flexible, waiting carries option value, and that value peaks exactly when the forecast is least certain, which inverts the usual instinct to lock early when the market feels frightening. Write down, before the negotiation opens, the rate level at which you would sign and the level at which you would wait.
When to change mode. The trigger is a ratio rather than a level. Air to ocean cost per kilo moves in a wide band and compresses hard when ocean spikes, so the same air quote can be indefensible one quarter and obvious the next without the air rate moving. A rate forecast contributes the probability of crossing that ratio, and the inventory consequences belong elsewhere (FF1).
One filter applies to all four. If the decision is the same at the tenth percentile and at the ninetieth, stop refining the forecast. A large share of freight forecasting effort produces precision that no decision consumes.
Where this stops
The largest moves in freight rates over the last five years came from events no statistical model was going to anticipate. Red Sea diversions from late 2023. Draught restrictions cutting Panama Canal transits through the 2023 to 2024 drought. The congestion of 2021 and 2022. Port labour disputes announced weeks ahead with unknowable outcomes. A model fitted across those episodes will reproduce them enthusiastically and predict none of the next ones, so the interval you publish should be built from a residual distribution that includes them while the point forecast leaves them alone. Holding both positions at once is the honest treatment.
A correlation problem sits underneath the plan as well. Freight spikes arrive with demand booms, because the same restocking cycle drives both, so a plan pairing a cautious freight assumption with an optimistic volume assumption has assumed the two are independent, and the errors add when they move together. Run one scenario where high volume and high rate land in the same quarter and see what it does to cash.
The blocker for most teams is more mundane. Reconstructing your own paid rate history at lane and week granularity is often impossible, because the freight invoice lives in accounts payable coded to a cost centre, with no container count and no link to the shipment, which leaves nothing to measure and nothing to score.
Start with the ten lanes carrying most of your freight spend. Pull twelve months of paid all-in cost per container from the invoices, line each one up against the published index for that lane and week, and compute the standard deviation of the residual. That single number tells you how much protection index-linking would actually have given you, and it sizes the freight reserve your next plan should carry.