MindForge Toolkit
Testing and review
Testing and review is the step before go-live: the builder checks their own work, and someone independent checks it again. It is one of the seventeen areas in the MindForge AI Risk Management Toolkit, the Singapore industry's practices for the AIRG. MindForge is the practices; the AIRG is the expectations and standards. I wrote the AIRG, and this guide sets out the practice and the expectation it meets.
What the AIRG expects
The AIRG expects evaluation and testing proportionate to materiality, across a range of plausible conditions, and a pre-deployment review by parties not involved in building the system, with formal independent validation for high-materiality use cases. The outcome it wants is evidence you can interrogate that the system is good enough for the task before it reaches customers or production data. For how this fits the lifecycle, see the AIRG in practice.
What MindForge says to do
Builders self-check first
Before anyone else looks, builders run AI risk self-checks that supplement, not replace, the firm's existing testing. They define objectives, design test cases to probe guardrail effectiveness, test performance against the risk-related metrics, and run fairness assessment, sensitivity analysis, sub-population analysis, error analysis, and stress testing. For generative and agentic systems, they test specifically for failure modes such as data leakage, toxicity, and hallucination, using data as specific to the firm and use case as possible. The results are documented as a precondition for deployment.
Then an independent AI-specific review
After the build and the self-checks, the firm runs an AI-specific review, sized to materiality, to confirm residual risks are identified and the governance was followed. It is most effective when carried out by someone not involved in developing or operating the use case, and completed before deployment. Bought AI is reviewed to the same standard as built AI, with compensatory testing where the firm cannot see inside the vendor's model. Findings, limitations, and conditions for use go to the approval body before sign-off.
In practice
What good looks like. A documented builder self-check covering the real failure modes, then an independent review by someone with no stake in the system shipping, who can and sometimes does send it back. A clear definition of what good enough means for the task, tested on firm-specific data, with bought AI held to the same bar.
Evidence to hold:
- The self-check results, including fairness, sub-population, error and stress testing, and the failure-mode tests for generative systems.
- The independent review report, its findings and conditions, and who performed it.
- A record of a use case sent back or made conditional, and the compensatory testing done on any bought AI.
How banks do it
The MindForge Implementation Examples show a cross-functional review body that meets before deployment: use-case owners present, risk specialists assess each risk type, and the body accepts, requests mitigations for, or rejects the use case. Systems built before the body existed are re-validated against current standards, and third-party technologies are flagged early in the outsourcing process.
My take
Committees are fine. The question is whether someone independent can say no, and does. And if you do not know what good enough for the task looks like, more testing will not tell you.
A review body that has approved everything it ever saw is a turnstile, not a gate, and testing without a definition of good enough just produces numbers. Fix the standard to the task first. (From my book, AI Risk Management for Directors.)
Work with me
I train and advise financial institutions on AI testing and independent review that actually bites, as part of the wider AIRG programme. See the courses and workshops, read more on AI risk management, or get in touch.
A guide in The MindForge AI Risk Management Toolkit, area by area. See also the AIRG, MindForge and CRI mapping.