Quaintitative

MindForge Toolkit

Use-case risk and materiality

How risky is this AI use case, and how much governance does it earn? That rating sets the weight of every control downstream, and it is one of the seventeen areas in the MindForge AI Risk Management Toolkit, the Singapore industry's practices for the AIRG. MindForge is the practices; the AIRG is the expectations and standards. This guide sets out the practice, the AIRG expectation it meets, and the evidence it produces. I wrote the AIRG, and the materiality rating is the number I would trust least until I had probed it.

What the AIRG expects

The AIRG expects a firm to rate each AI use on its materiality, minimally across impact, complexity and reliance, applied consistently, with the depth of controls following the rating, and material systems subject to independent pre-deployment review. The rating is how proportionality works: effort follows risk. For how it drives the rest, see the AIRG in practice and the AIRG explained.

What MindForge says to do

The Toolkit turns the rating into a working process in two passes, inherent then residual, with a review at the end.

Define materiality levels

Set clear, enterprise-wide tiers (commonly low, medium, high) on criteria that fit the firm: impact (financial, reputational, regulatory, recourse for affected people, use of personal data), complexity or novelty (multiple datasets, opaque architectures, agentic systems), and degree of reliance (how far AI substitutes for human judgement, and how little oversight sits over it). Tiers work only when they are unambiguous enough to be applied the same way across the firm.

Inherent, then residual

Assess the inherent risk early, from the use case's fundamental characteristics, with a flowchart of binary questions, a scored questionnaire, or both. For exploratory use cases, at least run a preliminary go/no-go and a rough tier before development, then a full assessment once the parameters are clear. After controls are in place, assess the residual risk in a setting representative of real use, with all guardrails applied, iterating controls and re-assessment until the residual risk is acceptable. A low-inherent use that meets its thresholds stays low; a high-inherent use such as credit decisioning stays high unless it reaches a particularly high standard.

Controls to match, and an independent review

Apply controls proportionate to the rating, at intensities that scale (sampling-based human review for a minor use, mandatory review of every output for a material one). Then an AI-specific review before deployment, conducted by a party not involved in building the use case, confirming the risks, the rating, and the mitigations. The review's depth and independence rise with materiality: a questionnaire and peer review for low risk, independent validation and the reviewer re-running the tests for high risk.

In practice

What good looks like. Materiality tiers defined unambiguously and applied the same way across the firm, on what a system does rather than who sees it. An inherent assessment early, a residual assessment before go-live after controls, and controls whose intensity visibly tracks the tier. A pre-deployment review by someone independent of the build, deeper and more independent as the risk rises.

Evidence to hold:

  • The materiality methodology with its tier definitions and worked examples.
  • Completed inherent and residual assessments for a sample of use cases across tiers, with who signed them.
  • The independent pre-deployment review, and evidence its depth scaled with the rating.

How banks do it

In the MindForge Implementation Examples, Prudential runs a risk-based framework aligned to its AI ethics principles, with pre-deployment review by independent domain experts and post-deployment reviews whose frequency follows materiality. UOB assigns each model a materiality rating from a weighted scoring methodology, and the rating plus the model's intended use determine the depth and independence of review, with independent validation for high-materiality models used in regulated processes.

My take

Rate by how much it can hurt, not who sees it, not what it is called. And clarity beats more rubrics.

Audience is the easy proxy and the wrong one: a back-office model can decide who gets paid. And a tier that two teams read differently is not a rating, it is a colour. Define it so a high means the same thing on every desk. (From my book, AI Risk Management for Directors.)

Work with me

I train and advise financial institutions on building a materiality assessment that holds, and the rest of the AIRG programme. See the courses and workshops, read more on AI risk management, or get in touch.