MindForge Toolkit
Monitoring
Monitoring is what keeps an AI system fit for purpose after it goes live. It is one of the seventeen areas in the MindForge AI Risk Management Toolkit, the Singapore industry's practices for the AIRG. This guide sets out the MindForge practice, the AIRG expectation it meets, and the evidence it produces. I wrote the AIRG, and monitoring is where a system that passed every pre-deployment check can still drift out of tolerance without anyone noticing.
What the AIRG expects
The AIRG expects ongoing monitoring of AI in production, proportionate to risk materiality, with tiered thresholds set to catch deterioration before it becomes harm, not just to record it after. It expects enhanced controls for AI designed to update itself automatically. The expectation is an outcome: the firm can show it watches its live AI and acts when something moves. For how monitoring sits with the other lifecycle controls, see the AIRG in practice and the AIRG itself.
What MindForge says to do
The Toolkit treats the monitoring plan written before deployment as the thing you now have to actually run. It covers the model, the data under it, the system around it, and the humans on top.
Monitor against thresholds, and act on breaches
Track performance and risk metrics, the effectiveness of each guardrail, and the uptake, cost and return of the use case, each against a threshold beyond which it is outside tolerance. A standardised evaluation framework keeps metrics, thresholds and frequency consistent across the firm, customised only where a technology makes a metric impossible or irrelevant. A red, amber, green dashboard communicates at a glance whether a use case has breached a threshold, is near one, or is within tolerance. When a threshold breaks, a guardrail fails, or user feedback raises a concern, find the cause and act: retrain, adjust, redevelop, roll back to a previous version, or activate a kill switch. Log results, incidents and escalations so root-cause analysis and later audit are possible, and so the metrics and thresholds themselves can be improved.
Watch the data, not just the model
Monitor data quality, drift and third-party data risks after deployment, for both training data and data used in operation such as grounding content in a retrieval setup, at a depth proportionate to materiality. Set statistical drift indicators and act when they fall outside the thresholds your risk appetite sets. Keep applying the data-management practices you already run, including access, security and timely deletion.
Keep re-checking the system, and the people
The use-case team should periodically self-check whether the things that set the controls have changed - the risk materiality, the scope of use, the key risks - and escalate where they have. Repeat AI-specific reviews after deployment, proportionate to materiality and triggered by events: a major model change, a threshold breach, stakeholder feedback, or an external development. Check that the people operating and overseeing the system are still skilled and supported as teams turn over.
Oversight, feedback and security
Operationalise human oversight proportionate to materiality, with a sampling approach that fits: targeted high-frequency review for higher-risk uses, lower-frequency review for lower-risk ones, and threshold-based review that only pulls in a human when a condition is met. Give end users a way to query, contest or give feedback on AI decisions, especially where the impact is material; even a thumbs up or down can be tracked as a metric that alerts on a rising share of negatives. For generative and agentic systems, monitor for adversarial use such as prompt injection with input and output filters, proportionate to risk.
In practice
What good looks like. Live metrics against thresholds that have actually tripped an action, drift monitoring on the data as well as the model, periodic re-checks of materiality and scope, human oversight sized to the risk, a feedback channel for the people affected, and a log that lets you reconstruct what happened. The depth scales with materiality; the discipline does not.
Evidence to hold:
- The monitoring dashboard and the alert history: what tripped, what was done, and how fast.
- Drift reports on training and operational data, with the thresholds and the actions taken.
- Post-deployment review records, override and feedback logs, and the incident log with root-cause and remediation.
How banks do it
The MindForge Implementation Examples show the pattern: banks run continuous performance and data-quality monitoring of their live AI through dashboards, feed in user feedback built into the system, and route potential issues to human review, so monitoring drives action rather than sitting as a report nobody reads.
My take
Watch one model at a time and you will miss what actually hurts you, the same weak model sitting in three places.
Monitoring that only looks system by system misses the portfolio view, which is where the real exposure builds. A threshold that has never tripped an action is not monitoring; it is a screen-saver. (From my book, AI Risk Management for Directors.)
Work with me
I train and advise financial institutions on monitoring AI in production and the rest of the AIRG programme. See the courses and workshops, read more on AI risk management, or get in touch.
A guide in The MindForge AI Risk Management Toolkit, area by area. See also the AIRG, MindForge and CRI mapping.