Quaintitative

AI forecasting

Why a language model cannot forecast on its own

Text and time look like cousins. They are not, and the difference is the whole problem.

When large language models went mainstream, the quickest way to spot someone overselling was that they claimed a chatbot could forecast something: a stock, a demand curve, next quarter's numbers. Anyone who said that had not understood how the model works. Things have moved on. If someone tells me today that they use AI for forecasting, I no longer dismiss it. I ask how. But there is still a firm line between a language model doing language and a forecasting model doing time, even when the forecasting model borrows the transformer architecture.

A language model is trained on text. Forecasting lives in the world of time series. Both are sequences, so they look related. They are not, and here is why.

Order is not a feature of the data. It is the data.

Take a table of data and shuffle the rows: the meaning survives, the same records in a different order. Reorder the parts of a sentence and the point usually still holds. Now take a time series and shuffle it. The signal is gone. Monday's sales followed by Thursday's followed by Tuesday's is not a rearranged story, it is noise. Sequence carries the meaning in text too, but it matters more in a time series, because the signal there is usually weaker to begin with.

A series has structure a language model has no notion of

Decades before anyone trained a neural network, statisticians read any series as a stack of parts: a level (where it sits), a trend (where it is drifting), seasonality (what repeats, the daily, weekly and yearly rhythms), and noise (what is left). A language model was never built to separate a trend from a wobble, because sentences do not have levels, trends and seasons.

A series is also restless. Words change meaning slowly; the sentence you read today parses as it would have a decade ago. Prices, demand, traffic and temperature do the opposite, and sometimes fast. The technical term is non-stationary, and together with noise it is a large part of why forecasting is hard: you are predicting a system that keeps changing underneath you.

And context is domain-specific. A stock series, a hospital-admissions series and an electricity-demand series share the word "time" and almost nothing else. Their noise, their seasonality and the way they drift are all different. Text has one grammar. Time has as many grammars as there are domains that produce it.

The confidence trap

This is what makes the "just use a chatbot" claim dangerous rather than merely wrong. Ask a language model to forecast when it has no way to actually do it, and it will not stop and say it cannot. It fills the gap the way it fills every gap, with fluent, plausible text: a number, a direction, a confident little paragraph about why. You will not be able to tell the good from the bad by looking, and often not even with judgment.

None of this makes language models useless for forecasting. The last few years have produced careful ways to make them work. But every one of them starts by taking the time problem seriously instead of wishing it away. The approaches that work are built around sequence, non-stationarity and domain-specific noise, rather than ignoring all three and hoping the architecture sorts it out. If you want to tell whether someone is selling you something real, that is the test: do they take the difference between text and time seriously, or not?

Read more

Next, the methods that do forecast: the forecasting ladder, and time-series foundation models. For the overview, see AI forecasting. This draws on the AI Agents for Forecasting primer.

Subscribe for updates