AI Supervision
How regulators assess AI
A supervisor does not assess a model. You assess whether a firm has managed it. That sounds like a small distinction, but how regulators assess AI turns on it entirely, and getting it wrong is the most common way supervision of AI goes soft. I supervised AI at the Monetary Authority of Singapore and wrote the AIRG, and this is the shift I come back to first.
How do we supervise something we cannot see inside? We did not build these models. We cannot read the weights. The firm knows its own system far better than we ever will.
All true. And none of it is new. So I ask back:
How do you supervise a credit model today? Have you ever rebuilt one to check it?
You have not, and you do not need to. You already supervise models no one in your team built. You do it through the firm's own validation, its limits, its reporting - laid out in a form you can interrogate and challenge. AI does not change the method. It raises the stakes on it.
The maths is the firm's job. The adequacy is yours
There is a reflex, when a technology feels unfamiliar, to think the supervisor must become a technologist to keep up. Resist it. If your test for supervising AI is whether your people can out-build the firm's data scientists, you have set a test you will always fail, and the wrong one.
Your object is not the model. It is the management of the model. Did the firm find this system and rate how much it matters. Did someone competent and independent of the builders check it. Is there a human who can actually intervene when it goes wrong. Is anyone watching it now that it is live. Is someone accountable for the outcome by name. Every one of those is a question about the firm's system, and you can judge all of them without reading a single weight.
This is also why a technologist alone does not make a good AI supervisor, and why a good model-risk or operational-risk supervisor is closer than they fear. The skills that transfer are the supervisory ones - reading evidence, spotting the gap between the document and the practice, knowing when a confident answer is covering a thin one. The AI specifics can be learned. The supervisory instinct is the scarce part, and you already have it.
Evidence you can interrogate, not a demo
Here is where it goes wrong in the room. A firm wanting to look strong will show you the system working. A smooth interface, a clean example, a confident presenter. A demo is not evidence. It is the one path the firm chose to walk you down, and it tells you almost nothing about the ninety-nine it did not.
What you are owed instead is evidence you can interrogate. The inventory entry. The materiality rating and the reasoning behind it. The validation report from someone who did not build the thing. The test results, including on the cases where it failed. The monitoring that has run since go-live, and what tripped. The definition of what "good enough" means for this task, and the test that proves the system clears it.
That last one is the quiet heart of it. Before any AI, the firm should be able to tell you what a good answer and a bad answer look like for the job this system does. Fraud detection has to catch fraud without burying people in false alarms. A customer-facing tool has to fit the customer and not mislead. If the firm cannot say what good enough is for the task, no amount of model sophistication saves them, and no amount of your technical knowledge fills the gap. You are not assessing whether the AI is clever. You are assessing whether the firm can show you the job still gets done to standard, with the AI in the mix.
A demo is the one path the firm chose to walk you down.
Why AI raises the stakes on an old method
The method is old. The reason it has to be applied harder is that an AI model has three properties ordinary software does not, and each one widens the gap between what a firm claims and what is true.
It is uncertain: it gives a spread of answers, not one, so "it works" is never the whole story and you have to ask how often, and how badly, it does not. It is unexpected: nobody wrote its rules by hand, so it can find a route to the goal that no one designed and no one anticipated, which is why testing has to reach past the happy path into the edge cases. And it is opaque: the more capable the model, the harder it is to say why it did what it did, which is why reproducibility and documented validation matter more here than for a system whose logic you could just read.
None of that means you fear the thing. A model is a mathematical mapping from inputs to outputs. It does not understand, intend, or scheme. Treating it like a person, in either direction, is the fastest way to misread it - trust it like a colleague and you over-trust it, fear it like a villain and you go hunting for motives when the cause is a skewed training set. You hold it to account as what it is: a system a firm chose to deploy, and must therefore be able to account for.
So the shift is this. Stop trying to assess the model. Assess whether the firm has found it, rated it, controlled it in proportion to what it matters, and can show you the evidence. That is a judgement you are already equipped to make. The rest of the book is how to make it well.
For the supervisor
What to look for. You are assessing the management of the model, not the model. The firm owns the maths; you own the judgement on adequacy, and you can reach every part of it through evidence rather than code. Refuse the demo. Ask for what you can interrogate: the inventory entry, the materiality rating and its reasoning, independent validation, test results including failures, live monitoring, and above all a definition of what "good enough" means for the task and the test that proves it. The AI specifics can be learned; the supervisory instinct you already have.
Ask the firm:
- For this system, what does a good answer and a bad answer look like - and how do you test that it clears the bar?
- Show me the evidence, not the demo: the validation report, the test results on the hard cases, the monitoring since go-live.
- Who checked this who did not build it, and what could they have sent back?
- If this system is wrong tomorrow, who answers for the outcome, by name?
Work with me
I train regulators, supervisors, and public authorities on AI governance and risk management. See the courses and workshops, read more on AI risk management, or get in touch.