AI forecasting
Time-series foundation models: the forecasting frontier
The top rung is genuinely new. It is also unfinished, and often beaten by a fifty-year-old method.
This is the new and evolving part, and where the "maybe AI can forecast now" story has real legs. The details matter, though. Taking an off-the-shelf chatbot like Claude or ChatGPT and claiming it can forecast is still overselling. A time-series foundation model is a different thing.
The idea: build a single model, pre-trained on enormous cross-domain piles of time-series data, that can forecast a series it has never seen, few-shot or zero-shot, the way a large language model answers a question on a topic no one fine-tuned it on. The aim is train once, forecast anything.
The obstacle, and the different answers to it
The obstacle is the structure of time from the start. Language has a finite vocabulary, a fixed set of words. Patterns in numbers are infinite. So how do you feed a continuous series into a transformer built for discrete tokens? Teams have answered differently. Amazon's Chronos chops the number range into 4,096 bins, in effect inventing a vocabulary for time. Google's TimesFM treats chunks of the series as patches, like image tiles. Salesforce's MOIRAI uses several patch sizes for different frequencies.
The field has forked. One camp trains a time-series foundation model from scratch: Chronos-2, MOIRAI-2, Lag-Llama, Nixtla's TimeGPT. Another camp does not build a new model at all; it takes a frozen language model and teaches it to handle numbers, as Time-LLM reprograms it, or as LLMTime feeds the series in as a string of digits. Chronos and TimesFM are open downloads on Hugging Face; Nixtla's TimeGPT may cost something to try. All are zero-shot: hand them a series, get a forecast, no training run.
Real progress, not a finished one
It is real progress. It is also not solved, and may take a while yet. When foundation models top the newer leaderboards, it pays to look closely: one careful study found scores inflated because the test data had leaked into training, and another put the leading models against decades of real stock returns and found they did not beat ordinary baselines off the shelf. The only thing that worked there was retraining from scratch on financial data, at which point you have a specialised model, not a general one. The top rung is exciting and worth watching. It does not retire the rest of the ladder, and it is still often beaten by rung one.
Read more
Why this points towards an agent rather than a single model: why agents, not prompts. For the overview, see AI forecasting. This draws on the AI Agents for Forecasting primer.