AI testing, evaluation and assurance
AI testing, evaluation and assurance in finance
You cannot govern AI you have not tested. Evaluation and independent assurance are where the controls stop being paper.
Testing and evaluation mean checking that an AI model or system does what it is meant to, on your own data and for your own population, before it goes live and while it runs. Assurance is the independent confirmation that it does, the evidence a second line, a board or a supervisor can rely on. Under the MAS AIRG, this is not optional: evaluation and testing are a lifecycle control that has to cover the key failure modes, and for generative AI and agents it has to test whether the guardrails actually hold, not just that they exist. For high-materiality AI, independent validation is a firm expectation.
A confident model that has not been tested on your population is a liability dressed as an asset, and a vendor's assurance is not your assurance. Assurance only counts when it is independent and competent, on your own data, against your own definition of good enough. A firm that can produce its evaluation results, its testing approach and its independent sign-off on request is in a different position from one that can produce only a policy.
Evaluation is not testing
They sound like the same thing and they are not. Evaluation measures how good a model is and tracks that number over time. Testing deliberately tries to break it, and you fix what fails. You need both. A good evaluation score tells you the model works on the cases you expected; testing tells you what happens on the cases you did not. Red-teaming is the most deliberate form of testing: structured but creative, time-boxed, and more useful when someone independent does it rather than the team that built the thing.
What makes an evaluation result trustworthy
A number on its own means little. Before you rely on an evaluation result, ask five questions:
- Ground truth - where did the right answers come from, and who curated them?
- Independence of the set - is the test set genuinely separate from the training data, and did anyone check for leakage?
- The metric - does it measure the decision the model actually makes? Classifying yes or no, predicting a number, ranking, and detecting anomalies each need their own metric.
- Runs per case - was each case run once, or enough times to average out the randomness?
- Who evaluated - the developers, an independent function inside the firm, or an outside party?
The last question is where assurance comes in. The developer's own evaluation is a starting point, not a sign-off. For high-materiality AI, the AIRG expects validation to be independent of the people who built and run the system.
Generative AI and agents make it harder
With a classifier you can check the answer against a known-correct label. With generative AI there is often no single correct answer, so you fall back on fuzzier measures: whether the output is close in meaning, or you ask another model to grade it, which imports that model's own biases. And what you are scoring shifts as the model is updated, so last quarter's result does not carry over.
Agents are harder still. A strong benchmark score does not predict how an agent behaves once it is wired into your systems and can act. Capability is not reliability. So you score two things, not one: the outcome, meaning did the task get done, and the trajectory, meaning was the path sound, the right tools called with the right arguments and no unsafe step along the way. And you score them more than once, because an agent that succeeds four times in five is not something you can depend on. The honest number is the share of runs that succeed every single time. If you use a model to judge the agent, the judge needs testing too: judges flip on small formatting changes, and tend to score outputs from their own family of models higher.
Where to start
- The AIRG in practice - where evaluation and testing sit in the lifecycle controls.
- For validators, reviewers and assurance - the seat that owns independent validation.
- The AIRG explained - the guidelines these expectations come from.
- The complete guide - the whole picture in one place.
- My research - including published work on evaluating AI systems, and runtime governance for agentic AI.
I developed the AIRG while leading AI risk supervision at MAS, and now advise firms on AI testing, evaluation and assurance independently. I am publishing more on this; subscribe to get it, and future updates on the AIRG.