AI forecasting
Why agents, not prompts, for AI forecasting
Stop asking the model to be the forecaster. Let it run the forecasting.
A language model cannot forecast on its own, and no single forecasting method is reliably best; the right one depends on the series in front of you. So the question changes. Not "can the model forecast?" but "can the model use the tools that can?" That reframing is what an agent is for.
Forecasting was never one step
Every rung on the ladder assumes forecasting is a single step: feed in history, get out a prediction. Anyone who has actually forecast knows it is a loop. Look at the series, is it trending, seasonal, drifting? Pick an approach that suits it. Run it. Then the part people skip: check whether it worked, look at the errors, backtest, see where it failed, and revise. Prepare, predict, critique, refine, round and round until it is good enough.
That loop is not a language task. But orchestrating it, deciding what to look at, which tool to call, what the errors mean and whether to try again, can be framed as language tasks. The newest forecasting research works this way: stop asking the model to be the forecaster and let it run the forecasting. The frameworks with grand names are this loop, automated.
What an agent actually is
An agent is a language model with three additions. Tools: real functions it can call, so instead of imagining an ARIMA forecast it runs one and gets real numbers back. The model selects which tool and with what inputs; the tool does the computation the model cannot. A loop: the ability to take a result, inspect it, and decide the next step rather than answering in one pass. Memory: holding what it has already tried across the steps, so the loop accumulates. The model handles the orchestration; the correctness comes from the tools. A model without tools produces confident nonsense, and tools without a model are just a library. Together they become a system that picks a real computation, runs it, checks what came back, and works towards a better result.
Why this beats prompting, concretely
Ask a chatbot to forecast and you get one confident paragraph with a number in it, from a system that cannot compute the number and will not tell you so, and that you cannot check. Give an agent the same task and, set up with real tools, it inspects the series, notices it is non-stationary, selects a method that handles that, runs a real backtest, sees the recent error is too high, tries a second method, compares them, and hands you the better one with the evidence, all logged.
The agent is not cleverer. Every number came from a tool that did the maths, and every choice is one you can inspect. The confidence trap, where a made-up forecast and a real one look identical, goes away, because a real computation and a real backtest now sit behind the answer. That matters beyond accuracy: if you have to answer for a forecast, to a board, a client or a regulator, "the AI said so" is not defensible, while "this method, chosen for this reason, backtested with this score" is.
It also ages better. Foundation models will improve and new rungs will appear; an agent does not mind, because a better forecaster is just a better tool to call. The orchestration stays the same while the tools underneath get stronger. You are not betting on one model being right. You are building the system that picks the right model.
Read more
The tools that sit under the agent: the tools a forecasting agent needs. For the overview, see AI forecasting. This draws on the AI Agents for Forecasting primer.