Quaintitative

Generative AI and agents in finance

Agentic and generative AI for finance

How a language model produces an answer, where generative AI actually helps in finance, what changes when a model can act, and the checks that follow. Plain, and deliberately boring.

Financial institutions are being sold generative AI that will transform their work, and agents that will do the work for them. Some of it is useful, some is hype, and in finance the difference matters, because these systems can affect customers, records and money. I do not teach prompt tricks, which I call prompt and pray. I teach how the system works and the checks you need before you let it act. I supervised AI risk at the Monetary Authority of Singapore and wrote Singapore's AI risk management guidelines. This is the short version.

How the machinery works

A text model predicts the next token; it does not check what it produces against an authoritative record. So the same fluent style can carry a correct figure and an invented one, a hallucination, and retrieval (RAG) feeds it your documents but does not guarantee a right answer. The answer is generated, so you need evidence it is right for the task.

Where generative AI helps in finance

One application is really several tasks, and they do not all need a language model, so decide per feature and let the boring parts stay boring. A prompt is an attack surface, so critical restrictions belong outside the model. A chatbot is a system, not a model.

What changes when a model can act

An agent proposes actions and a harness executes the permitted ones, so one request can span many steps, each looking fine while the path is wrong. Give an agent only the tools the job needs. The unit of governance is no longer the output. It is the trajectory: the request, the tool calls, the approvals and retries, recorded so someone can reconstruct what happened and with whose authority.

The checks that follow

Build evals (representative cases, clear criteria, repeated runs), and enforce the critical rules in the application, not the model's instructions, because a rule you can talk it out of is not a control. Put a person before the actions that matter, and monitor with measures someone can act on, because an indicator you cannot act on is a statistic.

That is the short version. The full treatment, the training and retrieval, the prompt and the system around it, the agent loop, the evals, the runtime controls and the records, is in the book: a short primer, and the full Agentic and Generative AI for Finance.

Frequently asked questions

What is the difference between generative AI and an AI agent?

Generative AI produces content, text, images or other output, by predicting it. An AI agent uses a model to select steps and call tools to carry out a task, with surrounding software, the harness, that checks and executes the permitted actions. An agent can act; a generative model on its own only produces an answer.

Do large language models hallucinate, and does that matter in finance?

Yes. A language model generates the most likely next piece of text; it does not check the answer against an authoritative record, so the same fluent style can carry a correct figure and an invented one. In finance that matters because the output can affect customers, records and money, so you need evidence the answer is right for the task, not just that it reads well.

What is retrieval-augmented generation (RAG)?

A model's knowledge is frozen at its training cutoff, so for document tasks the application retrieves the relevant passages, puts them in the context, and asks for an answer based on them. Good retrieval is necessary but not sufficient: a strong model can misread a good source, and a citation is not proof, so inspect the retrieved material when an answer is wrong.

What checks do you need before letting an AI agent act in finance?

Give the agent only the tools the job needs and enforce the critical rules in the application, not in the model's instructions: a rule you can talk it out of is not a control. Build evals for answers and actions, keep a full record of the trajectory, put a person before the actions that matter with the authority to decide, and monitor with measures someone can act on.

Why is the trajectory, not the output, the unit of governance for agents?

An agent can take many steps, and each step can look fine while the path is wrong, for example using a stale price in a correct calculation to prepare the wrong order. A fluent final confirmation does not reveal that the system used the wrong account or skipped an approval, so you govern and record the whole trajectory: the request, the tool calls and results, the approvals and the retries, linked to the task.

Related, on using AI in finance

  • AI for investing - what AI agents can and cannot do in investing, the four patterns, and the gap between a demo and production.
  • AI for forecasting - why a language model cannot forecast on its own, the methods that do, and the agent that drives them.
  • Agentic AI risk management - governing AI that acts, in depth.

Work with me

I train financial institutions and their teams on using generative AI and agents safely, from how they work to the checks before they act. See the courses and workshops, or get in touch.