Quaintitative

MindForge Toolkit

Third-party onboarding

Onboarding is the point where a firm brings in AI it did not build, and so the point to catch the risk that AI carries with it. It is one of the seventeen areas in the MindForge AI Risk Management Toolkit, the Singapore industry's practices for the AIRG. MindForge is the practices; the AIRG is the expectations and standards. I wrote the AIRG, and since most of the AI a firm uses it buys rather than builds, this is where a lot of the real exposure sits.

What the AIRG expects

The AIRG expects onboarding, development and deployment controls for third-party AI to be adequate, including compensatory testing to address the information gap left by what a vendor will not show you (AIRG 4.11). For how third-party AI sits across the framework, see the AIRG in practice, and the AIRG overview.

What MindForge says to do

The Toolkit starts from the firm's existing procurement practices and layers AI-specific checks on top, proportionate to the use case's materiality, since bringing in third-party AI can raise that materiality.

Due diligence on the vendor

Confirm the third party has done the work: robust testing on the AI-specific metrics that matter for your use case, guardrails that function in your context, and monitoring where it applies, such as for a connected service. Plan for the communications you will need when the vendor has an incident, makes a change, or has an outage.

Your own tests, on your own data

Where proportionate to materiality, run your own tests rather than relying on the vendor's. Vendors optimise to known benchmarks, so a representative, context-specific test on your own data tells you far more about real performance. Cover data risks such as bias and leakage, performance risks such as hallucination, security risks such as adversarial attacks, and legal risks such as IP violations. For generative AI, red-team the system for jailbreaks and misuse before it goes live, because standard testing misses the context-dependent, multi-turn vulnerabilities of these systems.

In practice

What good looks like. Due diligence proportionate to materiality that confirms the vendor tested, guarded and monitors the model; your own tests on your own data and population, not just the vendor's benchmark claims; red-teaming for generative tools; and a plan for vendor incidents, changes and outages.

Evidence to hold:

  • Vendor due-diligence records, covering the vendor's testing, guardrails and monitoring.
  • Your own test results on your own data, including red-team logs for generative systems.
  • The vendor incident, change and outage communication plan.

How banks do it

In the MindForge Implementation Examples, the pattern is to go beyond vendor assurances: firms run their own representative tests using their own data, and red-team generative tools for jailbreaks and misuse, before onboarding, rather than taking benchmark claims at face value.

My take

A vendor can prove their model is excellent on benchmarks and it tells you almost nothing about how it does on your task. Test it on your own data, or you have not really onboarded it.

Buying the model does not outsource the result you owe your customers and your regulator. The onboarding test on your own data is how you take that responsibility back. (From my AI Risk Management from First Principles primer.)

Work with me

I train and advise financial institutions on onboarding third-party AI and the rest of the AIRG programme. See the courses and workshops, read more on AI risk management, or get in touch.