Quaintitative

AI forecasting

The tools a forecasting agent needs

An agent is only as good as its tools. For forecasting there are three kinds, and the one people skip is the one that matters most.

An agent is only as good as the tools you give it. For forecasting they map onto the three things a forecaster actually does: get the data ready, make the prediction, and check whether it was any good. Miss one and the whole thing falls over.

Data tools: getting the series ready

Before anything forecasts, the series has to be fit to forecast on, and in forecasting the preparation is where most of the real skill hides. Data tools fetch the series, prices, demand, sensor readings, from wherever it lives. They clean it, filling gaps and handling outliers. They engineer features, the step that still wins competitions: lagged values, rolling averages, and the calendar facts (day of week, month, holiday) that carry seasonality. And they build the windows, slicing a long series into the input-and-target pairs a model learns from.

One data operation to never get wrong: splitting the data for testing without letting the future leak into the past. In most machine learning you split at random. In forecasting you cannot, because a random split lets the model peek at the future and every score afterwards is a lie. Time splits respect the order of time. A data tool that gets this wrong poisons everything downstream.

Model tools: making the prediction

These are the rungs of the ladder, each wrapped so the agent can call it: give it a series, get back a forecast. A classical tool (ARIMA, ETS, Prophet). A machine-learning tool (gradient-boosted trees on engineered features). A deep-learning tool. A foundation-model tool (Chronos or TimesFM). The point of wrapping them is that the agent does not need to be any of these models; it needs to know they exist, roughly what each is good for, and how to call one and read the result. It can run three and compare, or start simple and climb only if simple fails. When the agent reports a number, that number came out of one of these tools, not out of the language model. And if a model fails, you get an error, not a made-up answer. Failing loudly is a virtue here.

Evaluation tools: checking whether it worked

This is the one people skip, and skipping it is the difference between a forecasting system and a pattern generator. An evaluation tool answers the only question that matters: is this forecast any good? It runs a backtest, replaying history as if forecasting forward through it, so you see how the model would have done on data it did not train on. It computes the error with metrics built for forecasting, MASE, sMAPE and the like, not the generic ones. And it compares against a baseline, because a forecast is only impressive relative to the dumb alternative. Beating "tomorrow will look like today" is a low bar that is not easy to clear, and a surprising number of expensive models never clear it, sold by people who never checked.

Evaluation is also what lets the loop actually loop. The critique-and-refine cycle runs on it: without a real measure of how badly a forecast did, "try again" means nothing. It is the feedback, and it is what makes the confidence trap go away, because it forces every forecast to prove itself against reality before you trust it.

The three together

Data tools prepare the series and split it honestly. Model tools produce candidate forecasts from different rungs. Evaluation tools backtest them and pick the winner, or send the agent back to try again. The agent is the part that moves between them: prepare, predict, check, refine. That is a forecasting system.

Read more

For the overview, see AI forecasting; for why this approach beats a prompt, why agents, not prompts. The working code, notebooks and agent patterns are in the AI Agents for Forecasting primer and the fuller ebook it leads into.

Subscribe for updates