Quaintitative

Resource

AI risk management in finance:
a curated resource list

The sources worth knowing for managing the risk of AI in financial services: supervisory guidance, standards, model risk management, tooling, and research. Curated, and kept current.

AI is landing in banks, insurers, and asset managers, and in the authorities that supervise them. The discipline for managing it is not new. A firm that already runs model risk management, third-party risk, and technology risk has most of the muscle, and AI stretches each of them rather than replacing them. This list gathers the primary sources that matter, read through a finance lens throughout. It covers what people variously call AI risk management, AI governance, and responsible AI, all narrowed to financial services.

It is deliberately narrow. Where a horizontal standard or tool is included, it is because it is routinely used in finance.

On this page: Regulators · International bodies and standards · Industry frameworks · Model risk management · Tools · Research and reading

Financial regulators and supervisory guidance

Supervisory guidance from the authorities that regulate financial institutions, organised by jurisdiction. These are the texts a firm is actually held to.

Singapore (MAS)

United States

United Kingdom

European Union and member states

Canada (OSFI)

Asia-Pacific and other

International bodies and standards

Cross-border financial bodies and the horizontal standards that finance firms map onto.

Global financial standard-setters and cross-border bodies

Horizontal standards and frameworks applied to finance

Horizontal laws and conventions

Industry and consortium frameworks

Frameworks and toolkits from industry consortia and professional bodies, built for financial services.

Consortium and industry-body frameworks

Professional-body certifications and training

Trade-association and survey research

Professional-services frameworks

  • Deloitte Trustworthy AI framework (Deloitte) - Advisory firm's responsible-AI framework spanning transparency, accountability, fairness, privacy, safety, and robustness across the AI lifecycle; firm-branded, with financial-services applications.
  • KPMG Trusted AI framework (KPMG, 2025) - Advisory firm's framework for embedding governance across the AI and AI-agent lifecycle, used in its financial-services AI assurance work; firm-branded.

Model risk management foundations

The supervisory and practitioner base that AI risk management in finance builds on, and how it extends to machine learning and generative AI.

Supervisory foundations

Extending MRM to AI, ML and generative AI

Validation practice

Books and long-form

Tools and open source

Very little AI risk tooling is built for finance specifically; the main finance-native option, the open-source Veritas toolkit, sits under Industry and consortium frameworks. What follows is the general-purpose toolbox financial institutions actually use, grouped by the control each one serves:

  • Fairness and bias tools run the fair-lending and underwriting testing behind ECOA and Regulation B in the US and the FEAT principles in Singapore.
  • Explainability tools produce the adverse-action reasons credit decisions require, and feed independent model validation.
  • Validation, monitoring and drift tools run the ongoing monitoring and validation that SR 11-7, SS1/23 and OSFI E-23 expect.
  • LLM and GenAI evaluation and AI security tools cover generative-AI model risk and red-teaming.
  • Risk catalogues and databases help populate an AI inventory and risk register.

Open source unless marked commercial. A few canonical references are flagged where no longer actively maintained.

Fairness and bias

Explainability

  • SHAP (Scott Lundberg and community) - Game-theoretic Shapley-value method to attribute a model's output to its input features, widely used to generate adverse-action reasons and support model validation.
  • LIME (Marco Tulio Ribeiro) - Local interpretable model-agnostic explanations for individual predictions on tabular, text and image data; a standard for per-decision reason codes.
  • InterpretML (Microsoft) - Unified framework offering glassbox models (Explainable Boosting Machines) and blackbox explainers, valued where regulators expect inherently interpretable credit models.
  • Alibi (Seldon) - Library of explanation algorithms including counterfactuals and anchors, helpful for actionable recourse and adverse-action narratives.
  • Captum (Meta / PyTorch) - Model interpretability library for PyTorch (attributions, integrated gradients), relevant to validating deep-learning models in finance.
  • DALEX (ModelOriented) - Model-agnostic explanation framework in R and Python emphasising model exploration and fairness checks, common in actuarial and credit-risk workflows.
  • Responsible AI Toolbox (Microsoft) - Dashboards and libraries combining error analysis, interpretability, fairness and counterfactuals for structured model assessment and governance documentation.
  • Amazon SageMaker Clarify (AWS, commercial) - Managed service for bias detection and feature-attribution explainability across the ML lifecycle, used by regulated firms already on AWS.

Validation, monitoring and drift

  • Evidently AI (Evidently) - Open-source framework to evaluate, test and monitor ML and LLM systems with 100+ metrics for data and prediction drift, central to production model-risk monitoring.
  • Deepchecks (Deepchecks) - Holistic testing library for data and model validation from research to production, supporting the continuous validation expected under model-risk frameworks.
  • whylogs (WhyLabs) - Open-source data-logging library that produces statistical profiles for data-quality and drift tracking; WhyLabs offers a commercial observability platform on top.
  • NannyML (NannyML) - Open-source library that estimates post-deployment model performance without labels and links drift to performance impact, addressing silent model degradation in production.
  • Giskard (Giskard-AI) - Open-source evaluation, testing and red-teaming library for ML and LLM/agent systems, including an automated vulnerability scanner for bias, robustness and LLM risks.
  • Great Expectations (GX) - Widely used data-validation framework that enforces expectations on pipelines, foundational for the data-quality controls underpinning model inputs.
  • Arize Phoenix (Arize AI) - Open-source AI observability and evaluation platform (tracing, evals, drift), with a commercial Arize counterpart for managed production monitoring.
  • AI Verify (AI Verify Foundation / IMDA Singapore) - Open-source AI governance testing framework running technical tests and process checks against internationally recognised AI principles.

LLM and GenAI evaluation

  • DeepEval (Confident AI) - Open-source LLM evaluation framework with metrics for hallucination, relevancy and more in a unit-test style, useful for GenAI model-risk sign-off.
  • Ragas (exploding gradients) - Evaluation library focused on RAG pipelines (faithfulness, context precision and recall), relevant where finance GenAI must stay grounded in source documents.
  • promptfoo (promptfoo) - Declarative CLI/CI tool for LLM evals and red-teaming across providers, enabling repeatable regression and vulnerability testing of GenAI features.
  • OpenAI Evals (OpenAI) - Framework and open registry of benchmarks for evaluating LLMs and LLM systems, usable to build private domain-specific evals.
  • LangSmith SDK (LangChain) - Open-source client SDK for the LangSmith platform (tracing, datasets, evaluation); the SDK is open, the hosted platform commercial.
  • TruLens (TruEra / Snowflake) - Open-source library to evaluate and track LLM apps and agents with feedback functions (groundedness, relevance, toxicity), supporting GenAI assurance.
  • Project Moonshot (AI Verify Foundation / IMDA Singapore) - Open-source toolkit combining benchmarking and red-teaming to evaluate the safety and reliability of large language models and LLM applications.

AI security and red-teaming

Risk catalogues and databases

  • MIT AI Risk Repository (MIT FutureTech) - Structured, regularly updated database of over 1,000 documented AI risks with a causal and domain taxonomy, usable as a starting catalogue for enterprise AI risk registers.
  • AI Incident Database (Responsible AI Collaborative) - Searchable index of real-world AI harms and near-harms, a source of loss scenarios and precedent for operational-risk and model-risk analysis.
  • AI Vulnerability Database / AVID (AVID) - Open knowledge base of failure modes for general-purpose AI systems with reproducible evidence and a taxonomy library, supporting structured GenAI vulnerability tracking.
  • MITRE ATLAS (MITRE) - An ATT&CK-style knowledge base of real-world tactics and techniques against ML systems, useful for threat modelling AI in financial infrastructure.

Research, reports, and reading

Landmark reports, papers, books, courses, and people worth following.

Reports and research

Academic and practitioner papers

Books

Courses and training

Blogs, newsletters and people

  • Eugene Yan (eugeneyan.com) - Applied ML practitioner writing substantively on LLM evaluation, eval design and production ML, directly useful for AI risk and model testing.
  • AI Snake Oil (Narayanan and Kapoor) - Princeton-led newsletter critically analysing AI capabilities, evaluation and overclaiming, including predictive AI in consequential domains.

Reference hubs

Adjacent curated lists worth knowing. None is finance-specific, which is the gap this one fills.

More on this site