MindForge Toolkit
Third-party onboarding
Onboarding is the point where a firm brings in AI it did not build, and so the point to catch the risk that AI carries with it. It is one of the seventeen areas in the MindForge AI Risk Management Toolkit, the Singapore industry's practices for the AIRG. MindForge is the practices; the AIRG is the expectations and standards. I wrote the AIRG, and since most of the AI a firm uses it buys rather than builds, this is where a lot of the real exposure sits.
What the AIRG expects
The AIRG expects onboarding, development and deployment controls for third-party AI to be adequate, including compensatory testing to address the information gap left by what a vendor will not show you (AIRG 4.11). For how third-party AI sits across the framework, see the AIRG in practice, and the AIRG overview.
What MindForge says to do
The Toolkit starts from the firm's existing procurement practices and layers AI-specific checks on top, proportionate to the use case's materiality, since bringing in third-party AI can raise that materiality.
Due diligence on the vendor
Confirm the third party has done the work: robust testing on the AI-specific metrics that matter for your use case, guardrails that function in your context, and monitoring where it applies, such as for a connected service. Plan for the communications you will need when the vendor has an incident, makes a change, or has an outage.
Your own tests, on your own data
Where proportionate to materiality, run your own tests rather than relying on the vendor's. Vendors optimise to known benchmarks, so a representative, context-specific test on your own data tells you far more about real performance. Cover data risks such as bias and leakage, performance risks such as hallucination, security risks such as adversarial attacks, and legal risks such as IP violations. For generative AI, red-team the system for jailbreaks and misuse before it goes live, because standard testing misses the context-dependent, multi-turn vulnerabilities of these systems.
In practice
What good looks like. Due diligence proportionate to materiality that confirms the vendor tested, guarded and monitors the model; your own tests on your own data and population, not just the vendor's benchmark claims; red-teaming for generative tools; and a plan for vendor incidents, changes and outages.
Evidence to hold:
- Vendor due-diligence records, covering the vendor's testing, guardrails and monitoring.
- Your own test results on your own data, including red-team logs for generative systems.
- The vendor incident, change and outage communication plan.
How banks do it
In the MindForge Implementation Examples, the pattern is to go beyond vendor assurances: firms run their own representative tests using their own data, and red-team generative tools for jailbreaks and misuse, before onboarding, rather than taking benchmark claims at face value.
My take
A vendor can prove their model is excellent on benchmarks and it tells you almost nothing about how it does on your task. Test it on your own data, or you have not really onboarded it.
Buying the model does not outsource the result you owe your customers and your regulator. The onboarding test on your own data is how you take that responsibility back. (From my AI Risk Management from First Principles primer.)
Work with me
I train and advise financial institutions on onboarding third-party AI and the rest of the AIRG programme. See the courses and workshops, read more on AI risk management, or get in touch.
A guide in The MindForge AI Risk Management Toolkit, area by area. See also the AIRG, MindForge and CRI mapping.