Frameworks compared
AI model validation in finance
What SR 11-7, SS1/23, OSFI E-23, the EU AI Act, the MAS AIRG, NIST and ISO 42001 each require for building, testing and validating AI models, and where they converge and diverge.
Cross-cutting synthesis
This is the area where the model risk management tradition does most of the work, and where the frameworks converge most tightly. The common backbone is the one SR 11-7 set out in 2011: a model is validated through conceptual soundness, testing before use, and an independent challenge of the whole thing, under the discipline of effective challenge - a critical review by people with the competence, independence and authority to probe a model and force change. Every broad AI-risk framework restates it. The MAS AIRG folds it into development standards plus independent validation or peer review calibrated to risk materiality; SS1/23 makes independent validation its own principle; OSFI E-23 requires reviews that confirm a model is fit for purpose; the EU AI Act requires testing and a conformity assessment before a high-risk system reaches the market; and the horizontal standards - NIST's Measure function and ISO/IEC 42001's pre-deployment gate - say the same for any sector.
Two adaptations recur for AI specifically:
- Validation scales to materiality. The AIRG calibrates the depth of validation to a model's risk materiality, and the risk-based tiering in SS1/23 and OSFI E-23 does the same: a high-impact credit or capital model attracts full independent validation, while a low-risk tool may need only peer review. The more complex and more opaque the model, the deeper and more independent the challenge it earns.
- The techniques stretch for AI. More weight falls on testing, robustness and explainability; validation must confirm stable performance on new data, guard against overfitting, and verify outputs can be reproduced in production. Generative AI shifts to outcome-based and human evaluation, benchmarking and red-teaming, because a probabilistic generator cannot be back-tested like a scorecard.
The divergence is about instrument and prescriptiveness - from the EU's formal conformity assessment to the UK's and US's principles applied to all models - not about whether rigorous, independent, pre-deployment validation is required.
What each framework says
MAS AIRG - validation is a lifecycle control, calibrated to risk materiality.
- Algorithm and feature choices must be justified against simpler alternatives.
- Testing is proportionate and broad - out-of-sample, stability, sub-population, stress and adversarial methods, with overfitting guarded against.
- High-risk AI gets independent pre-deployment validation of conceptual soundness, data, fairness, explainability and limitations, plus a technology and cybersecurity review.
- Generative AI is tested harder for hallucination, data leakage and adversarial attacks.
Guidelines on AI Risk Management (MAS, 2025)
United States, SR 11-7 - three-element validation under effective challenge. Conceptual soundness, ongoing monitoring and outcomes analysis, reviewed by parties with the competence, independence and authority to probe a model and force change. The 2026 revision carries this forward and puts generative and agentic AI out of scope. SR 11-7 / OCC 2011-12 (Federal Reserve / OCC)
United Kingdom, PRA SS1/23 - independent review of every model. The technique must be conceptually sound and supported by research or accepted practice, and risk-based tiering prioritises validation effort. The principles are technology-neutral, applying to all models regardless of technology, including AI and machine learning. SS1/23: Model risk management principles for banks (Bank of England / PRA)
United Kingdom, PRA / FCA DP5/22 - AI's complexity may need more supervisory focus on testing, validation and explainability than traditional models require. DP5/22: Artificial Intelligence and Machine Learning (Bank of England)
Canada, OSFI E-23 - reviews scaled to the risk rating. Model reviews confirm a model is properly specified, working as intended and fit for purpose, triggered by new models, changes, performance breaches or scheduled review, and outputs must be replicable in production. Guideline E-23: Model Risk Management (OSFI)
European Union, AI Act - test and conformity-assess before market. High-risk systems run an iterative lifecycle risk process, must meet accuracy, robustness and cybersecurity standards, and must pass a conformity assessment before deployment. Regulation (EU) 2024/1689, the Artificial Intelligence Act
ECB, internal-models regime - an independent validation function, and machine learning raises the bar. Validation assures model quality including back-testing; the 2025 guide adds a machine-learning chapter treating ML as a driver of higher complexity, materiality and validation expectations. Guide to internal models (ECB Banking Supervision, 2025)
NIST AI RMF - the Measure function: test, evaluate, verify, validate. Metrics, validity, reliability and robustness, measured independently where possible and documented so results can be challenged. AI Risk Management Framework 1.0 (NIST)
ISO/IEC 42001 - a documented pre-deployment gate. Verification and validation across the AI lifecycle, backed by internal audit of the management system. ISO/IEC 42001:2023, AI management system
MAS Veritas (FEAT) - the principles turned into verifiable assessment steps. A repeatable way to test and document AI and data-analytics systems before and after deployment. Veritas Document 3: FEAT Principles Assessment Methodology (MAS)
Hong Kong, HKMA - rigorous validation and testing before production use, to confirm the accuracy and appropriateness of AI models. High-level Principles on Artificial Intelligence (HKMA)
Hong Kong, SFC - validate and monitor generative-AI models. Adequate validation and performance assessment, with ongoing review, applied in a risk-based way to the use case. Circular on the Use of Generative AI Language Models (SFC)
IOSCO - testing and monitoring before deployment, so systems behave as expected under varied conditions, with senior management accountable for the lifecycle. Artificial Intelligence in Capital Markets (IOSCO, 2025)
IAIS - regular, documented performance assessment that accounts for known model and data limitations. Application Paper on the Supervision of Artificial Intelligence (IAIS)
AIR and Google Cloud - outcome-based and human evaluation for generative AI. Grounding, benchmarks and human evaluation with continuous monitoring, robust testing and human-in-the-loop, because accuracy metrics alone are insufficient for generative outputs. Generative AI Risk Management in Financial Institutions (AIR and Google Cloud)
Comparison
| Framework | Position on development, testing and validation |
|---|---|
| MAS AIRG | Development standards plus independent validation or peer review calibrated to materiality; pre-deployment checks; stricter generative-AI testing |
| SR 11-7 / SR 26-2 (US) | Three-element validation (conceptual soundness, ongoing monitoring, outcomes analysis) under effective challenge by independent parties |
| PRA SS1/23 (UK) | Conceptually sound technique; independent review of all models; risk-based tiering prioritises validation |
| PRA / FCA DP5/22 (UK) | Flags a greater need for testing, validation and explainability for AI |
| OSFI E-23 (Canada) | Reviews confirm properly specified, working as intended, fit for purpose; frequency by risk rating; outputs replicable in production |
| EU AI Act | Testing before market; accuracy, robustness and cybersecurity standards; mandatory conformity assessment for high-risk |
| ECB (internal models) | Independent internal validation function; machine learning raises complexity, materiality and validation expectations |
| NIST AI RMF | Measure function: test, evaluate, verify, validate; documented and challengeable |
| ISO/IEC 42001 | Documented pre-deployment gate; verification and validation across the lifecycle; internal audit |
| MAS Veritas | FEAT operationalised into verifiable assessment steps for testing and documentation |
| HKMA | Rigorous validation and testing before production use |
| SFC (HK) | Adequate validation, performance assessment and ongoing review of generative-AI models |
| IOSCO | Thorough testing and monitoring before deployment; behaviour verified under varied conditions |
| IAIS | Regular, documented performance assessment accounting for model and data limitations |
| AIR / Google Cloud | Grounding and outcome-based evaluation; robust testing, human-in-the-loop, benchmarks plus human evaluation |
FAQ
What is the common backbone of model validation, and where does it come from? SR 11-7's three core elements - conceptual soundness, ongoing monitoring, and outcomes analysis - plus effective challenge by independent, competent, authorised reviewers. Later frameworks restate and extend it rather than replace it, so a firm with a mature model risk management function already has most of the muscle for AI validation.
Does validation have to be independent, or is peer review enough? It depends on materiality. SS1/23 requires independent review of all models; OSFI E-23 allows internal reviewers or objective third parties with depth scaled to the risk rating; and the MAS AIRG accepts independent validation or peer review calibrated to risk materiality. Higher-risk AI attracts genuinely independent challenge; lower-risk tools may be peer-reviewed.
How does the EU AI Act's approach differ from SR 11-7's? The EU AI Act imposes a formal, documented conformity assessment and binding accuracy, robustness and cybersecurity requirements for high-risk systems before they reach the market. SR 11-7 and SS1/23 instead apply technology-neutral validation principles to all models through supervisory expectation. The EU route is prescriptive and gated; the US and UK route is principles-based and continuous.
Why does generative AI need different testing? A generative model cannot be back-tested like a scorecard. The guidance shifts to outcome-based and human evaluation, grounding and retrieval quality, benchmark suites, guardrail and red-team testing, and human-in-the-loop review, because the risks - hallucination, prompt injection, data leakage - are not captured by conventional accuracy statistics.
Does using machine learning automatically raise validation expectations? In the ECB's internal-models regime, yes: it treats machine learning as a driver of complexity and therefore materiality, which brings higher expectations towards reporting and internal validation. Other frameworks reach the same place through risk-based tiering: more complex, more opaque models earn deeper and more independent validation.
Read the sources
Every framework quoted above is in the full AI risk management in finance resource list. For the Singapore picture, see the AIRG explained and the AIRG compared to the EU AI Act, NIST and ISO 42001.
Work with me
I train banks, insurers, and supervisors on turning model validation and AI risk management into a working system, grounded in the AIRG and the frameworks above. See the courses and workshops, or get in touch.