Quaintitative

AI forecasting

The forecasting ladder: classical models to deep learning

Every rung trades complexity for capability, and some capability for risk. Climbing it is not automatic progress.

There is a ladder of ways to forecast, and reaching for the newest, largest model because it is newest and largest, skipping what would have worked better, faster and cheaper, is usually a mistake. Here are the rungs, and the forty-year record of what actually wins.

Rung one: the classical models that still win

For decades, forecasting belonged to statisticians, and they built methods that took the structure of time seriously from the start. The main one was ARIMA, and its name is a summary of what matters: AutoRegressive (today depends on yesterday's values), Integrated (take differences until the series stops drifting, the classical answer to non-stationarity), and Moving Average (today also depends on yesterday's errors). Around it grew a family, each aimed at one property of time: SARIMA for seasonality, VAR for several related series at once, Holt-Winters for trend and season separately, and, in finance, ARCH and GARCH for volatility that is itself a moving target.

Two things about this rung get forgotten. It is explainable by design: because the models are built out of trend, seasonality and cycles, you can say why they forecast what they did. And it still wins: fifty-year-old methods regularly beat the latest neural network, which is why a baseline is not a nicety but the only way to know whether a more complex model actually improved anything. R's forecast package and Python's statsforecast cover the whole family; Meta's Prophet offers the classical-style additive decomposition of level, trend and season.

Rungs two and three: machine learning and deep learning

When machine learning became powerful, people pointed it at forecasting: random forests, support vector machines, gradient-boosted trees. The trick was often to turn time into tabular columns, the day of week, the month, and so on. It worked sometimes. One kind earned a lasting place: gradient-boosted trees fed with hand-built features can be genuinely excellent, not because they understand time but because the features are relevant (lightgbm and xgboost).

Then the deep-learning wave arrived from vision and language. Sequence models such as RNNs and LSTMs were built to remember what matters from earlier in a sequence; probabilistic forecasters like DeepAR became popular; and after the transformer, a flood of models followed, N-BEATS, the Temporal Fusion Transformer, Informer, Autoformer, PatchTST. The results were mixed. Sometimes a deep model was excellent; sometimes a simple classical model beat it in a fraction of the time. Deep learning did buy one genuinely new thing: it could handle multimodal series, numbers and news, prices and events, which classical models could not. That is the rung my own research lived on, fusing prices with news over evolving company networks, and it produced strong but specialised models, trained for one task rather than general-purpose (neuralforecast and darts).

The lesson of forty years

The Makridakis Competitions have pitted forecasting methods against each other for four decades, and the pattern is stubborn. The early ones were won by classical methods. M4 (2020) was won by a hybrid, exponential smoothing married to an LSTM, not pure deep learning. M5 (2022) was won by gradient-boosted trees with feature engineering, not deep learning at all. M6 (2025), which pushed into finance, saw most teams underperform simple benchmarks.

Read across forty years: the wins credited to deep learning are usually wins from feature engineering, ensembling or hybrids. The fancy model is rarely the reason. And where newer models do top a leaderboard, it pays to look closer. One careful study found benchmark scores inflated because test data had leaked into training; another put leading foundation models against decades of real stock returns and found they did not beat ordinary baselines off the shelf, until they were retrained from scratch on financial data, at which point you have built a specialised model rather than a general one.

So the ladder is real, but climbing it is not automatic progress. The top rung is genuinely new and often beaten by rung one. If no single model is reliably best, and the right choice depends on the series in front of you, then something has to inspect the data, pick the rung, run the model and check whether it worked. That is where an agent comes in.

Read more

Next: time-series foundation models (the top rung), and why agents, not prompts. For the overview, see AI forecasting. This draws on the AI Agents for Forecasting primer.

Subscribe for updates