AI Supervision
The AI supervision assessment checklist
This is the working reference a supervisor can take into a firm: the areas to assess in a firm's AI risk management, what to look for in each, and the one question to put, mapped to the AIRG. I supervised AI at the Monetary Authority of Singapore and wrote the AIRG. The paragraph references and the areas are the MAS'. The "look for" and the "ask" are my own interpretations, not regulatory guidance, and are deliberately compressed; the rest of the book carries the reasoning. Use it as a prompt sheet, not a tick-box. The order follows the AIRG's architecture: oversight, the risk-management system, the lifecycle controls, and capability.
Oversight
1. Board responsibilities (AIRG 2.5). Look for: challenge, not noting - a board that questioned and sometimes pushed back, with material AI risk in an appetite that can actually be breached, and competency adequate to what is deployed. Ask: Show me where the board changed or stopped an AI proposal, and the AI limits in the appetite someone could breach.
2. Senior management (AIRG 2.6). Look for: the framework turned into practice, resourced in proportion to the estate, with breaches escalated in time. Ask: How many people run AI risk oversight, against how many systems, and how was that sized?
3. Cross-functional committee (AIRG 2.4). Look for: where overall AI risk is material, a committee judged by the decisions that bite, not its terms of reference. Ask: Show me a decision this committee made that changed or stopped something.
4. Org-wide integration (AIRG 2.3). Look for: AI folded into existing model, operational, technology, and third-party risk frameworks, not sitting beside untouched ones; watch the waiver log. Ask: Which existing risk policies did you update for AI, and how often do AI deployments get waivers?
5. Risk culture (AIRG 2.2). Look for: responsible-use principles turned into practical red lines and a real channel to raise concerns, not a poster. Ask: What has a staff member actually been stopped from doing with AI under your policy?
The risk-management system
6. Identification (AIRG 3.2). Look for: a net, not a list - a firm-wide definition, a control function accountable for it, and real detection of shadow and embedded AI. Ask: How do you find the AI that is not on the inventory, and when did you last catch one you had missed?
7. Inventory (AIRG 3.4). Look for: one reconciled inventory with populated metadata, real linkages, a cadence that matches deployment, and a decommissioning log. Ask: Reconcile the AI inventory against your IT, vendor, and model registers - what is in one and not the other?
8. Materiality (AIRG 3.8). Look for: rating on what a system does (impact, complexity, reliance), calibrated so a "high" means the same everywhere, with an honest distribution. Ask: Take me to three low-rated systems - why low, what moves each up, and who signed?
Lifecycle controls
9. Data management (AIRG 4.5). Look for: data representative of production and traceable to source, not "fine"; due diligence on external data. Ask: How did you confirm the training data represents the population this system meets in production?
10. Transparency and explainability (AIRG 4.6). Look for: explanations a customer can act on and a reviewer can challenge, matched to materiality - not a technique applied for show. Ask: Show me a real explanation this system gave a customer or reviewer. Could they act on it?
11. Fairness (AIRG 4.8). Look for: a chosen fairness definition suited to the use, proxy variables tested, pre-deployment sub-group testing, and any accuracy trade-off made in the open. Ask: What fairness definition did you choose and why, which proxies did you test, and what did fairness cost in accuracy?
12. Human oversight (AIRG 4.10). Look for: effective oversight - real override authority, logged and analysed overrides, sustainable workloads, active handling of automation bias. Ask: Show me the overrides on this system and what you learned; what would a rate near zero tell you?
13. Third-party AI (AIRG 4.11). Look for: bought AI held to the same standard, with compensatory testing on the firm's own data and a usable exit. Ask: What testing have you done on your own data, as opposed to what the vendor told you - and could you actually leave this vendor?
14. Model selection (AIRG 4.12). Look for: a simple baseline actually built and beaten, with complexity consciously justified. Ask: What simpler model did you build as a baseline, and what did the more complex one buy you over it?
15. Evaluation and testing (AIRG 4.14). Look for: a stated "good enough" threshold, no train-test leakage, real edge-case and adversarial testing, and results that include the failures. Ask: What does good enough mean for this task, and show me the results on the cases where it did worst.
16. Technology and cyber (AIRG 4.16). Look for: the system secured against ordinary technology risks plus AI-specific ones like data poisoning. Ask: What AI-specific attacks did you test against, and what did the security review conclude before go-live?
17. Reproducibility and auditability (AIRG 4.17). Look for: development documented well enough for an independent reviewer to rebuild it, with versioned data, code, and configuration. Ask: Can you reproduce this model from your records - the data version, code, and configuration?
18. Pre-deployment validation (AIRG 4.18). Look for: genuinely independent validation for material systems, with authority to send a system back and a record that it sometimes has. Ask: Show me a system your validation function sent back or made conditional, and whether those conditions closed.
19. Monitoring (AIRG 4.23). Look for: tiered thresholds that have actually tripped actions, segmented so a failing sub-group is not hidden in a healthy average. Ask: Show me the alert history - what tripped a critical threshold, what was done, and how fast?
20. Change management (AIRG 4.25). Look for: a clear material-versus-minor line, retrains sent back through review, and controlled emergency and unauthorised changes. Ask: When this system was last retrained or upgraded, did it go back through review, or just ship?
21. Pilots and PoCs (AIRG 4.3, fn 14). Look for: time-boxed pilots with end dates that have not silently passed, and controlled use of production data. Ask: Give me your active pilots with their end dates - which have passed, and is production data in them?
22. Auto-updating AI (AIRG 4.25c). Look for: enhanced controls - strict limits on what can change on its own, and a rollback that has been tested, not just configured. Ask: What can change automatically on this system, and when did you last test the rollback?
Capability
23. Competence and capacity (AIRG, capability). Look for: skill across all three lines - validators who can actually challenge a model - and headcount sized to the estate; treat a gap as a finding. Ask: Can your validators actually challenge a model or only process forms, and what is your plan for the capability gap you already know you have?
Two closing reminders
The grid lists areas in order, but risk does not arrive in order. Two cross-cutting moves matter more than any single row. First, point your scrutiny where risk hides - the low rating, the inventory gap, the committee that has never said no - not where the firm wants to show you its best work. Second, hold everything to proportionality: the firm's controls must track its ratings, and its ratings must survive your scrutiny. Get those two right and the twenty-three rows above become a map you can actually use, rather than a checklist you tour.
Work with me
I train regulators, supervisors, and public authorities on AI governance and risk management. See the courses and workshops, read more on AI risk management, or get in touch.