AI Supervision
The AI inventory and materiality rating
If you only had an hour with a firm, you would spend it on the AI inventory and the materiality rating. Everything else in AI risk management hangs off two things: whether the firm knows where its AI is, and whether it has rated each use honestly. I supervised AI at the Monetary Authority of Singapore and wrote the AIRG, and this is the part of supervision I would protect above all the rest. Supervisors often treat the inventory as a formality, a list to confirm exists, and the rating as the firm's internal business. Both are where the real supervision is.
The firm has an AI inventory and a materiality methodology. We have seen both documents. Is that not the box ticked?
The documents existing is the least interesting fact about them. The question that matters is whether the inventory catches the AI the firm would rather not log, and whether the ratings survive a second look, because those two are where risk actually goes missing. A control only reaches the AI on the list, and every control downstream is dialled by the rating. Get these two right and the rest of the framework has something to stand on. Get them wrong and the best lifecycle controls in the world are protecting the wrong systems, or none.
Identification: the net, not the list
Before an inventory there is identification, how the firm decides what counts as AI and goes onto the list in the first place. The AIRG calls consistent identification across all business and functional areas a critical prerequisite, and puts a control function in charge of it firm-wide rather than leaving each unit to decide for itself. That last point matters to you: if identification is devolved, every unit draws the boundary where it suits them, and the firm-wide picture fragments.
So assess the net, not the list. Start with the definition: is there a clear, firm-wide criterion for what counts as AI, and who arbitrates the borderline cases. A firm with no stable definition cannot have a reliable inventory, because the thing being inventoried keeps changing shape. Then test the hard edges, which is where shadow AI lives. AI arrives in a firm without knocking: a feature switched on inside a SaaS tool already licensed, a vendor model embedded in a product no one logged as AI, a model that grew out of a spreadsheet, a team using a public chatbot on personal accounts because the sanctioned route was slow. Ask how procurement surfaces AI bought inside something else. Ask how the firm detects unauthorised use. Ask when the inventory last gained an entry someone had missed, and how it was caught. A firm that can describe only its list, and not how it finds what is off the list, is telling you the net is thin.
Watch consistency across units, geographies, and legal entities too. The same kind of system should be identified the same way everywhere. Where one business line logs a tool and another running the same tool does not, you have found both an identification failure and, usually, the softer unit.
The one inventory, and what it has to carry
The AIRG wants a single, accurate, up-to-date inventory, because one firm-wide inventory is what lets risk be aggregated and controls applied consistently. Several inventories that do not reconcile are no inventory at all.
Judge it on four things. Completeness, which you test by reconciling it against other registers, the IT asset list, the vendor register, the model inventory, and seeing what appears in one and not the AI inventory. Metadata, because an entry that is just a name is useless: each record should carry the materiality rating, the business and technical owner, the use case and its authorised boundaries, data sources and sensitivity, deployment status, dependencies, validation status, and performance where relevant. An inventory where mandatory fields are half-populated is an inventory no one actually uses. Linkages, because a system does not stand alone, the record should connect to the data feeding it, the applications depending on it, the models it leans on, so that when something upstream changes, the firm can see what downstream is affected. And currency, because an inventory updated annually against an estate that changes monthly is stale by design; look at the update triggers and cadence against the pace of deployment.
One quiet check reveals discipline: decommissioning. Ask to see the log of retired systems. A firm that only ever adds to its inventory and never removes anything is not maintaining it; it is accreting. The retirement log tells you whether the inventory is a living control or a graveyard that also happens to list the living.
The cheapest way to make risk disappear is to rate it low.
The rating: on what it does, and whether it is honest
Now the most consequential number in the firm's whole AI programme. The materiality rating sets how hard every downstream control has to work, so it is both the thing you most need to be sound and the cheapest place in the system to make risk quietly disappear.
The AIRG requires a methodology that rates each use at minimum across Impact, Complexity, and Reliance, the harm if it fails, how novel and opaque it is, how much the firm leans on it with how little human oversight and how little reversibility. The first thing to check is that the firm rates on these, on what the system does, and not on the easy proxy of who sees it. Audience is not materiality: a system no customer sees can decide which claims get paid, and a customer-facing system can be a harmless chatbot. If a firm's methodology sorts by audience, it will mis-rate, and you should expect to find high-consequence systems hiding in low-audience tiers.
Then check calibration, whether a "high" means the same thing in every unit. Look for scoring rubrics with worked examples, and for a mechanism that keeps ratings consistent across business lines, because without one, materiality drifts into whatever each team can justify. And check reassessment: a rating is not permanent, and the methodology should name the triggers that force a re-rate, a change of use, a performance degradation, an incident, because a system rated low two years ago may have quietly become critical.
The sharpest single check is the one from where firms hide risk, and it belongs here in full. Do not audit the high-rated systems; they are being cared for. Sample the low-rated ones and pull the thread. Why is this low. What would move it up a tier. Who signed the rating, and were they independent of the people who benefit from it staying low. Then step back and read the distribution across the whole estate. If almost everything sits in the lower tiers, far more than a firm of that size and business should produce, that shape is itself a finding, either unusually little risky AI, or systematic rating-down. Make the firm show you which. The control function the AIRG puts in charge of materiality exists precisely so the rating is not set by whoever benefits from the answer; confirm that arbiter is real and that it has, at least once, overruled a unit's self-rating.
Why this is the part that matters
Spend your scrutiny here and the rest of your supervision gets easier, because a firm that genuinely knows where its AI is and has rated it honestly has a map you can rely on, and proportionality, your own and the firm's, finally works. Spend your scrutiny everywhere else and skip this, and you are inspecting controls without knowing whether they sit on the systems that need them. The inventory and the rating are not the boring preliminaries to the real work. They are the real work, and almost everything the firm might hide from you hides in the gap between the two.
For the supervisor
What to look for. Identification that works as a net, not a list: a clear firm-wide definition, a control function accountable for it, and real detection of shadow AI at the edges (SaaS features, embedded vendor models, unauthorised use); consistency across units. One reconciled inventory, not several: tested for completeness against the IT, vendor, and model registers, with populated metadata, real linkages to upstream data and downstream systems, a cadence that matches deployment, and a decommissioning log that proves it is maintained. And a materiality rating that turns on what the system does (impact, complexity, reliance), not who sees it: calibrated so a "high" means the same everywhere, reassessed on defined triggers, and arbitrated by a control function. Sample the low-rated, read the portfolio distribution for an unnatural lean to the lower tiers, and make the firm account for the shape.
Ask the firm:
- What is your definition of AI, who arbitrates the borderline cases, and how do you find the AI that is not on the inventory?
- When did your inventory last gain an entry someone had missed, and how was it caught?
- Reconcile your AI inventory against your IT asset, vendor, and model registers for me. What appears in one and not the AI list?
- Show me your decommissioning log. What have you retired this year?
- Take me to three low-rated systems. Why low, what would move each up a tier, and who signed the rating?
- Show me your materiality distribution across the estate. Why does it lean the way it does, and has the control function ever overruled a unit's self-rating?
Work with me
I train regulators, supervisors, and public authorities on AI governance and risk management. See the courses and workshops, read more on AI risk management, or get in touch.