Quaintitative

For Risk & Compliance

Challenging AI before go-live

Before an AI system ships, the business builds it. In the second line, you do not build alongside them. You challenge what they built, and for the material systems you are the gate it has to pass. Challenging and validating AI before go-live is model-risk validation stretched to AI, which is familiar ground: I wrote the AIRG and led the thematic review of how banks actually manage AI model risk, so this is how I would run the gate.

Faced with the list of development controls - data, selection, explainability, fairness, testing, security, documentation, validation - second lines do one of two things. They tour all of them shallowly, or they fixate on one and miss the rest.

There are so many lifecycle controls. Do we really have to validate every one of them for every system?

No, and trying would be the proportionality mistake from a few chapters back. So the question sharpens:

Which of these controls, if it is weak, makes all the others unreliable - and which one, challenged properly, tells you most about the whole build?

The controls are not a flat checklist to tick. They chain. Bad data makes a good model unsound; an untested model makes every claim about it a guess; and nothing the business did before go-live counts for much if no one independent of the builders checked it. That last independent check is you. Challenging a model you did not build, against its data, its testing, and its documentation, is model-risk validation stretched to AI. Extend the discipline you have.

Data: press representativeness and lineage

A model is its training data wearing a function. If the data was wrong, the model is wrong in ways no later control fully fixes, so this is where you challenge first.

Press representativeness hardest: does the training data actually match the population and the conditions the system will meet in production, or was it whatever happened to be on hand? A credit model trained on a narrow slice of borrowers will fail quietly on everyone outside it. Then quality - completeness, accuracy, timeliness - with the plain red flags: heavy missing data in the features that matter, labels of unknown reliability, external data taken on trust from a vendor with no due diligence. Then lineage: can the business trace a given output back through the transformations to the source, which is what makes an incident investigation possible at all. You are not redoing their data science. You are refusing to accept "the data is fine" as anything other than an assertion with evidence behind it, or without.

Selection: make them justify the complexity

The AIRG asks developers to justify and document why they chose the model they did, especially when they reach for a more complex one. Use that. It is one of your cleaner challenges.

Ask whether a simple baseline was built and beaten. The honest process builds a plain model first, measures it, and only moves to something more complex if the gain is real and worth the cost in opacity and risk. Where there is no baseline, there is usually no discipline, just a reach for the most powerful thing on the shelf, with the complexity accepted by default rather than chosen. A marginal accuracy gain bought with a large loss of explainability is a trade someone should have made on purpose, and should be able to defend to you. If they cannot say what the simpler option would have cost, they did not really choose, and you can send them back to find out.

Transparency: test a real explanation

The AIRG ties transparency to materiality - more for high-stakes decisions, lighter for low-risk ones - and judges it by whether the explanation lets someone make an informed decision, not by whether a technique was run.

So test two things. For customers: is the disclosure explicit that AI is being used, in plain words, and does an adverse decision come with a reason the person could act on - what would have to change for a different outcome - rather than "computer says no" dressed in legal language. For the firm's own people: do the humans meant to oversee the system get explanations good enough to actually challenge it. A SHAP chart no reviewer understands is not explainability, it is decoration that passes an audit. Ask to see a real explanation that went to a real customer or reviewer, and ask whether they could act on it.

Fairness: defined, tested, and traded off in the open

The AIRG asks the firm to define what it considers fair for its context, and to find and mitigate harmful bias, proportionate to materiality. The trap is a fairness policy that never touches a model.

Check that the business has actually chosen a fairness definition and metric suited to the use case, and can say why, because the different definitions genuinely conflict and no model satisfies all of them at once. Check that it has identified not just the obvious protected attributes but the proxies that stand in for them, since a model can discriminate through a postcode without ever seeing a protected characteristic. Check that bias was tested before deployment, across sub-populations, with sample sizes large enough to mean something. And check that where fairness cost some accuracy, that trade was made in the open and signed off, not buried. For a lending or pricing model this is not a soft control. It is where your conduct and regulatory exposure concentrates, which is where your challenge has to be sharpest.

If you have never sent a system back, you are not a gate. You are a turnstile.

Evaluation and testing: the control you press hardest

If you challenge one control hard, challenge this one, because it sits upstream of almost everything else you care about. The AIRG wants evaluation and testing proportionate to materiality, across a range of plausible conditions from the ordinary to the edge cases. Done well, it tells you how uncertain the system is, where it breaks, and whether it clears the bar for the task. Done badly, every other claim about the system is a guess.

The failures are specific, and you should know them cold. Test data that leaked from training, so the reported performance is a flattering illusion. Testing only on the happy path, with no edge cases, no stress, no underrepresented groups. For systems where it applies, no adversarial or red-team testing - no probing for prompt injection or jailbreaks in a generative system, no evasion testing in a classifier. And underneath it all, no stated threshold: no definition of what "good enough" means for the task, so the test cannot pass or fail, only produce numbers. That last one is the quiet heart of the book - if the business cannot tell you what a good answer and a bad one look like for the task, the testing proves nothing, however much of it there is. Ask for the results on the cases where the system did worst. A team that only has results where it did well has not finished testing.

Security, and the record that lets you check anything

Two controls here do quiet, load-bearing work for you. The first is technology and cyber: an AI system is still a system, and it has to be secure against the ordinary technology risks plus the AI-specific ones like data poisoning. The second is reproducibility and auditability, which is the control that makes your validation possible at all. The AIRG wants development documented well enough that a competent, independent reviewer - you - could rebuild the work: the data versions, the code, the training procedure, the configuration, archived so a production model can be traced to the exact run that produced it. Without it, "we validated it" cannot be verified, an incident cannot be investigated, and a drifted model cannot be compared to what it used to be. A team that cannot reproduce its own model has handed you nothing to check.

The gate is you

Everything above is the build. This is the gate, and the gate is you. The AIRG requires that an AI use case be reviewed before deployment by parties not involved in building it, and that high-materiality systems undergo formal independent validation. Not a peer glancing over a colleague's work. Independence with teeth, which means your teeth.

So hold three things on yourself. That you are genuinely independent of the builders, with no stake in the system shipping. That your review has the authority and the depth to matter, so you can send a system back or attach conditions proportionate to the risk, and have done so. And that when you approve with conditions, you track them and confirm they close, rather than noting them and moving on. The hardest question here is the one you ask yourself: when did I last send a system back, and what did I demand? A validation function that has approved everything it ever saw is not a gate. It is a turnstile, and the business walks through it at will.

For the second line

What to own. The challenge, and the gate. Press the development controls as a chain, proportionate to materiality: data representative of production and traceable to source; a model choice justified against a baseline actually built; transparency a customer and a reviewer can act on; fairness with a chosen definition, proxies tested, trade-offs in the open; and above all evaluation with a stated "good enough" threshold, no train-test leakage, real edge-case and adversarial testing, and results that include the failures. Insist on reproducibility, because without it you cannot validate at all. And own the gate for material systems: genuinely independent, with the authority to send a system back, conditions tracked to closure, and a record that you have used it.

Ask the first line:

  • How did you confirm the training data represents the population this system meets in production, and can you trace an output back to its source?
  • What simpler model did you build as a baseline, and what did the more complex one buy you over it?
  • Show me a real explanation this system gave a customer or a reviewer. Could they act on it?
  • For a lending or pricing model: what fairness definition did you choose and why, which proxies did you test, and what did fairness cost in accuracy?
  • What does "good enough" mean for this task, what threshold did you set, and show me the results on the cases where it did worst.
  • Can you reproduce this model from your records - the data version, code, and configuration?

Work with me

I train risk and compliance teams on turning AI risk management into a working system. See the courses and workshops, read more on AI risk management, or get in touch.