AI Supervision
Proportionality in AI supervision
Proportionality is the one idea that makes supervising AI possible at all. It is also the one most often mistaken for a way out. Firms reach for the word when they want to do less. Supervisors flinch from it for the same reason. Both have it backwards. I supervised AI at the Monetary Authority of Singapore and wrote the AIRG, and proportionality done right is risk-based AI supervision in a sentence: effort follows risk, for the firm's controls and for your scrutiny alike.
As the number of AI systems climbs into the hundreds, how is a supervisor ever meant to keep up? We cannot look at everything.
No, you cannot, and you should stop wishing you could. So the question turns:
Why would you want to? Would you spend the same hours on the tool that drafts meeting minutes as on the model that decides which claims get paid?
Of course not. The same instinct that makes that obvious for you is the instinct the firm is meant to apply to its own controls. Proportionality is not permission to skimp. It is the rule that effort follows risk - for the firm in how it controls, and for you in how you scrutinise. Used honestly it is the only way either of you scales. Used as cover it is the first thing a firm hides risk behind, which is why you have to be able to tell the two apart.
Two systems, same audience, worlds apart
Start with how risk gets sized, because a firm that sizes it badly makes everything downstream meaningless. The intuitive proxy is audience - personal, internal, customer-facing - and it is not enough. Take two systems no customer ever sees. One drafts meeting minutes. The other feeds the decision on which insurance claims get paid. Same audience. One is harmless; the other moves real money and can treat people unfairly. Sort by who sees the system and you will rate those two the same, which is wrong.
The AIRG rates on what the system does, across three dimensions, and they are worth holding in your head because you will test every rating against them. Impact: the harm if it fails - to customers, financially, operationally, to reputation. Complexity: how novel and opaque the thing is, how hard to explain, how tangled its data dependencies. Reliance: how much the firm leans on it, with how little human oversight, and how reversible the outcome is. A claims model scores high on all three; the minutes tool scores low. The rating is the firm's judgement, and your job is not to re-run it but to probe whether it holds - which is exactly the first place firms hide risk.
Proportionality done by the firm
Once a system is rated, proportionality should visibly change what happens to it. That is the thing to check: not that the firm has a materiality policy, but that the rating actually pulls different controls.
A high-materiality system should carry heavier controls down its whole life - formal independent validation before go-live, closer monitoring with tighter thresholds, more frequent review and re-validation on a schedule, documentation deep enough to reproduce the work. A low-materiality system should carry lighter controls, and that is correct, not lax. The failure is not that low systems get less. The failure is when the rating and the controls come apart: a system rated high that got the light-touch treatment anyway, or a whole estate where everything, high and low, gets the same thin once-over. Either way the rating has stopped doing its job.
So the test is linkage. Pick a high-rated system and a low-rated one and trace what each actually received. If a "high" bought nothing a "low" did not, the firm has a materiality policy on paper and flat controls in practice. That is a finding, and a common one.
Effort follows risk - or the rating means nothing.
Proportionality done by you
Now turn it on yourself, because the same rule is what lets you supervise an estate you could never inspect in full. You do not owe every AI system the same depth of attention, and spreading yourself evenly is not fairness, it is waste - thin everywhere, so thin where it matters. Match your scrutiny to materiality the same way you expect the firm to match its controls. The material systems get the deep look: the validation interrogated, the monitoring read, the owner questioned. The trivial ones get a light touch, or a check that the firm has rated them trivial for good reason. Your scarce hours go where the harm would be.
There is one move that keeps this from collapsing into the firm's own self-assessment. Spend a slice of your attention not on the high-rated systems but on the ratings themselves - the sample of low-rated ones you went to first. Supervise the firm's proportionality, not just its systems. If the firm's sizing is honest, you can lean on it and go deep only where it points. If its sizing is self-serving, you will find that by auditing the ratings, and then you cannot trust the map and must look wider. Proportionality for you is therefore conditional: it is as good as the firm's materiality process, so you earn the right to rely on it by testing that process first.
Why this is the tool, not the loophole
The word gets a bad name because firms deploy it to argue for less. But strip the motive out and proportionality is just the recognition that risk is not evenly spread, so effort should not be either. A regime that demanded the full treatment for every AI use would collapse under its own weight, and firms would either ignore it or drown, and the risky systems would get no more care than the trivial ones. That helps no one.
Proportionality done right concentrates care where harm lives. The supervisor's whole craft is to make sure "proportionate" means that, and not "less wherever we can get away with it." You do that by holding two things together: the firm's controls must track its ratings, and its ratings must survive your scrutiny. Get those two right and the number of systems can climb as high as it likes. Your method does not break, because it was never trying to look at everything in the first place.
For the supervisor
What to look for. Proportionality is the rule that effort follows risk - for the firm's controls and for your scrutiny alike. Check that the firm rates on what a system does, not who sees it (impact, complexity, reliance), and that the rating actually pulls different controls: trace a high-rated and a low-rated system and confirm the "high" bought something the "low" did not. Then apply the same rule to yourself - deep attention to the material systems, light touch to the trivial - but earn the right to rely on the firm's map by auditing a sample of its ratings first. Proportionality for you is only as trustworthy as the firm's materiality process.
Ask the firm:
- Show me a high-rated and a low-rated system side by side. What did each actually receive in validation, monitoring, and review?
- Does your rating turn on what the system decides, or on who sees it?
- What concretely changes when a system is rated one tier higher - which controls switch on?
- How do you keep ratings consistent across units, so a "high" on one desk means a "high" on the next?
Work with me
I train regulators, supervisors, and public authorities on AI governance and risk management. See the courses and workshops, read more on AI risk management, or get in touch.