Quaintitative

AI Supervision

Monitoring AI after deployment

A system that was fine at launch does not stay fine. The world moves, the data shifts, the firm retrains it, and the controls that fit last year go slack. Monitoring AI after deployment, from the supervisor's seat, is checking whether the firm notices. I supervised AI at the Monetary Authority of Singapore and wrote the AIRG, and this is where a lot of real-world AI failure actually happens, long after the launch everyone celebrated.

Once a system has passed validation and gone live, it is governed, is it not? The hard work is the approval.

Approval is a snapshot of one moment. So the question is: what happens to this system between now and the next review, who is watching it, who can stop it, and what happens when it is changed? The risk does not hold still after go-live. It drifts, and the controls that matter now are the ones that catch the drift: a human who can actually intervene, monitoring that is actually read, change management that sends altered systems back through review, and the discipline to retire what should be retired.

Human oversight: effective, not nominal

Firms love to point at the human in the loop. The AIRG does not accept the human as a control by itself, and neither should you. It wants oversight proportionate to materiality and, crucially, effective, which means the person can actually intervene, the system was designed from the start to let them, and the arrangement accounts for automation bias and decision fatigue.

Those last two are where nominal oversight dies. Automation bias: a reviewer who approves whatever the model suggests is not oversight, they are a rubber stamp with a pulse, and over time most unsupported humans drift there. Decision fatigue: a reviewer handed a thousand cases an hour cannot meaningfully check any of them, so the span of control is itself a risk indicator. So you look past the existence of a reviewer to the evidence the oversight bites. Do reviewers have genuine authority to override, and is overriding actually feasible rather than a buried, discouraging process. Are overrides logged, and does anyone analyse the pattern, because a near-zero override rate and a very high one both mean something is wrong. Are reviewers competent for the decisions they are checking, and is their workload sustainable. The single question: show me the overrides, and what you learned from them. No overrides, or no analysis of them, means the human in the loop is a decoration.

Monitoring: configured is not the same as watched

A live model decays. Data drift, where the inputs shift away from what the model was trained on; concept drift, where the relationship the model learned stops holding; plain performance decay. The AIRG wants comprehensive ongoing monitoring with tiered thresholds set to pre-empt deterioration, not just to record it after the fact.

The failure here is subtle and common: monitoring that is configured but not acted on. Dashboards exist, thresholds are set, alerts fire, and nothing happens, because the thresholds were tuned so loose they never trip, or so tight that everyone learned to ignore the noise. So you do not check whether monitoring exists. You check whether a breach has ever forced an action. Pull the alert history: what tripped, what was done, how long it took. Look for the tiered levels the AIRG asks for, informational, warning, critical, and confirm the critical ones actually escalate to someone who can act. And check that monitoring is segmented, because an aggregate accuracy that looks fine can hide a model that has quietly failed for one customer group. A monitoring system that has never triggered a response is either watching nothing or being ignored, and you need to know which.

A threshold that has never tripped an action is a decoration, not a control.

Change management: a retrain is a new model

Here is the gap firms fall into most. A system passes a rigorous validation, goes live, and is then changed, retrained on new data, repointed at a new use, swapped onto a new model underneath, without going back through anything like the review it first passed. The validation you relied on is now describing a system that no longer exists.

The AIRG wants controls that define what counts as a material change versus a minor one, and govern the material ones properly. So your first check is that definition: how does the firm distinguish a change that needs full re-review from one that does not, and who decides the ambiguous cases. Then the discipline: when this system was last retrained or upgraded, did it go back through review and re-validation proportionate to its materiality, or did it just ship. A retrain is a small new model, and the firm that treats it as a routine refresh is deploying unvalidated models continuously while believing it is governed. Watch two more things: emergency changes, which are legitimate but become a loophole if they are frequent and never reviewed after the fact; and unauthorised changes, which proper separation of duties between who requests, approves, and implements a change is meant to prevent.

And retirement, which is a control in its own right. The decommissioning discipline from the inventory shows up here as practice: when a system is retired, are the things that depended on it updated or shut off too, or does a dependent report quietly keep running on a model that no longer exists.

Pilots, and the systems that live in the gaps

Firms run pilots and proofs of concept under lighter controls, which is reasonable, you cannot demand full validation of an experiment. The AIRG allows this but expects clear policies governing the deviation, including limits on time and users. The supervisory check is almost comically simple and almost always finds something: do the pilots have end dates, and have any of them passed. A "pilot" that has quietly been in production for two years, serving real customers under experiment-grade controls, is a material system hiding behind a label. Ask for the list of active PoCs, check the dates, and ask whether production data is being used in them and under what protection. The gap between "pilot" and "production" is a favourite place for real systems to live ungoverned.

Auto-updating systems, and the plan for when it goes wrong

Two tail risks to close. First, systems designed to update themselves automatically. The AIRG wants enhanced controls here, strict justification, clear limits on what can change on its own, and the ability to roll back, because a system that retrains itself in production can drift past its validation without a human ever deciding to let it. Ask what can change automatically, what the guardrails are, and whether rollback has ever been tested rather than just configured.

Second, incident response. For high-risk AI the AIRG expects contingency plans with fallback options, and where kill-switches are used, activation protocols that are regularly tested. A kill-switch no one has ever pulled in a drill is a hope, not a control. Ask when the firm last rehearsed turning a material system off and running without it. The answer tells you whether the contingency plan is real or filed.

For the supervisor

What to look for. After go-live the risk moves, so check the controls that catch movement. Human oversight that is effective, not nominal - real override authority, logged and analysed overrides, sustainable reviewer workloads, and active handling of automation bias; a reviewer who approves everything is not a control. Monitoring that is watched, not just configured - tiered thresholds that have actually tripped actions, segmented so a failing sub-group is not hidden in a healthy average. Change management that treats a retrain as a new model and sends material changes back through review, with emergency and unauthorised changes controlled, and retirement that updates downstream dependencies. Pilots with end dates that have not silently passed. Enhanced controls and tested rollback for auto-updating systems. And incident plans - fallbacks and kill-switches - that have actually been rehearsed.

Ask the firm:

  • Show me the overrides on this system and what you learned from them. What is your override rate, and what would a rate near zero tell you?
  • Show me your monitoring alert history: what has tripped a critical threshold, what was done, and how fast?
  • How do you tell a material change from a minor one, and when this system was last retrained or upgraded, did it go back through review, or just ship?
  • Give me your list of active pilots and PoCs with their end dates. Which have passed, and is production data being used in them?
  • For systems that update automatically, what can change on its own, and when did you last test the rollback?
  • When did you last rehearse turning a material AI system off and running without it?

Work with me

I train regulators, supervisors, and public authorities on AI governance and risk management. See the courses and workshops, read more on AI risk management, or get in touch.