Quaintitative

For Risk & Compliance

Watching AI after go-live

A system that was fine at launch does not stay fine. The world moves, the data shifts, the business retrains it, and the controls that fit last year go slack. Monitoring AI after go-live is the second line's job of noticing before anyone outside does. None of it is new to you: it is the post-implementation monitoring and change control you already run for models and technology, stretched to cover systems that learn. I wrote the AIRG and led the thematic review of how banks actually manage AI model risk, and this is where a lot of the real failure turned up.

The reflex is to treat the go-live approval as the finish line. The system got its sign-off; it is governed; on to the next one in the queue.

Once a system has passed validation and gone live, it is governed, is it not? The hard part was the approval.

Approval is a snapshot of one moment. So the question is:

What happens to this system between now and the next review - who is watching it, who can stop it, and what happens when the business changes it?

The risk does not hold still after go-live. It drifts, and the controls that matter now are the ones that catch the drift: a human who can actually intervene, monitoring that someone actually reads, change management that drags altered systems back through your review, and the discipline to retire what should be retired. This is where most real AI failure happens, long after the launch everyone celebrated.

Human oversight: effective, not nominal

The business loves to point at the human in the loop. Do not accept the human as a control by itself. What you need is oversight that is proportionate to materiality and, more to the point, effective - the person can actually intervene, the system was built from the start to let them, and the arrangement accounts for automation bias and decision fatigue.

Those last two are where nominal oversight dies. Automation bias: a reviewer who approves whatever the model suggests is not oversight, they are a rubber stamp with a pulse, and over time most unsupported humans drift there. Decision fatigue: a reviewer handed a thousand cases an hour cannot meaningfully check any of them, so the span of control is itself a risk indicator. So look past the existence of a reviewer to the evidence the oversight bites. Do reviewers have genuine authority to override, and is overriding actually feasible rather than a buried, discouraging process. Are overrides logged, and does anyone analyse the pattern, because a near-zero override rate and a very high one both tell you something is wrong. Are reviewers competent for the decisions they are checking, and is their workload sustainable. The question to put: show me the overrides, and what you learned from them. No overrides, or no analysis of them, and the human in the loop is a decoration you are relying on.

Monitoring: configured is not the same as watched

A live model decays. Data drift, where the inputs shift away from what the model was trained on; concept drift, where the relationship the model learned stops holding; plain performance decay. What you want is comprehensive ongoing monitoring with tiered thresholds set to pre-empt deterioration, not just to record it after the fact.

The failure here is subtle and common: monitoring that is configured but not acted on. Dashboards exist, thresholds are set, alerts fire, and nothing happens, because the thresholds were tuned so loose they never trip, or so tight that everyone learned to ignore the noise. So do not check whether monitoring exists. Check whether a breach has ever forced an action. Pull the alert history: what tripped, what was done, how long it took. Look for the tiered levels - informational, warning, critical - and confirm the critical ones actually escalate to someone who can act, and that you are on that path. And check that monitoring is segmented, because an aggregate accuracy that looks fine can hide a model that has quietly failed for one customer group. A monitoring system that has never triggered a response is either watching nothing or being ignored, and you need to know which before the supervisor asks you.

A threshold that has never tripped an action is a decoration, not a control.

Change management: a retrain is a new model

Here is the gap firms fall into most, and the one that will quietly invalidate your own work. A system passes a rigorous validation, goes live, and is then changed - retrained on new data, repointed at a new use, swapped onto a new model underneath - without going back through anything like the review it first passed. The validation you signed is now describing a system that no longer exists.

So own the change gate. Set the definition first: what counts as a material change that needs full re-review, versus a minor one, and who decides the ambiguous cases, which should be you. Then hold the discipline: when a system is retrained or upgraded, it comes back through review and re-validation proportionate to its materiality, rather than shipping on the quiet. A retrain is a small new model, and a first line that treats it as a routine refresh is deploying unvalidated models continuously while believing it is governed. Watch two more things. Emergency changes, which are legitimate but become a loophole if they are frequent and never reviewed after the fact. And unauthorised changes, which the separation of duties between who requests, approves, and implements a change is there to prevent, and that separation is yours to insist on.

And retirement, which is a control in its own right. The decommissioning discipline from the inventory chapter shows up here as practice: when a system is retired, the things that depended on it are updated or shut off too, so a dependent report does not quietly keep running on a model that no longer exists.

Pilots, and the systems that live in the gaps

The business runs pilots and proofs of concept under lighter controls, which is reasonable, because you cannot demand full validation of an experiment. But the deviation needs clear policy, including limits on time and users, and the check is almost comically simple and almost always finds something: do the pilots have end dates, and have any of them passed. A "pilot" that has quietly been in production for two years, serving real customers under experiment-grade controls, is a material system hiding behind a label, and it is your exposure. Keep the list of active PoCs, check the dates, and check whether production data is being used in them and under what protection. The gap between "pilot" and "production" is a favourite place for real systems to live ungoverned.

Auto-updating systems, and the plan for when it goes wrong

Two tail risks to close. First, systems designed to update themselves automatically. These need enhanced controls - strict justification, clear limits on what can change on its own, and the ability to roll back - because a system that retrains itself in production can drift past its validation without anyone deciding to let it. Insist on knowing what can change automatically, what the guardrails are, and whether rollback has been tested rather than just configured.

Second, incident response. For high-risk AI you want contingency plans with fallback options, and where kill-switches are used, activation protocols that are regularly tested. A kill-switch no one has ever pulled in a drill is a hope, not a control. Make sure the firm has rehearsed turning a material system off and running without it, because the first time you find out whether that works should not be during the incident.

For the second line

What to own. After go-live the risk moves, so own the controls that catch movement. Set and run a review cadence driven by materiality, not a calendar. Hold human oversight to effective, not nominal - real override authority, logged and analysed overrides, sustainable workloads, active handling of automation bias; a reviewer who approves everything is not a control you can rely on. Make monitoring watched, not just configured - tiered thresholds that have actually tripped actions, segmented so a failing sub-group is not hidden in a healthy average, with you on the escalation path. Own the change gate: a clear material-versus-minor line, retrains sent back through review, emergency and unauthorised changes controlled, retirement that updates dependencies. Keep pilots inside their end dates. Insist on tested rollback for auto-updating systems and a rehearsed kill-switch for the material ones.

Ask the first line:

  • Show me the overrides on this system and what you learned from them. What is the override rate, and what would a rate near zero tell us?
  • Show me the monitoring alert history: what has tripped a critical threshold, what was done, and how fast - and am I on that escalation path?
  • How do we tell a material change from a minor one, and when this system was last retrained or upgraded, did it come back through review, or just ship?
  • Give me the active pilots and PoCs with their end dates. Which have passed, and is production data being used in them?
  • For anything that updates automatically, what can change on its own, and when did we last test the rollback and rehearse the kill-switch?

Work with me

I train risk and compliance teams on turning AI risk management into a working system. See the courses and workshops, read more on AI risk management, or get in touch.