Quaintitative

For Risk & Compliance

Challenge, not build

The second line does not build the model. It does not sign it off as fine either. It challenges it. That is the whole job of risk and compliance with AI, and the two ways it goes wrong are forgetting the first half and forgetting the second. I wrote the AIRG and led the thematic review of how banks actually manage AI model risk, and the strong second lines all understood this; the weak ones had slid into one trap or the other.

How do we challenge something we cannot see inside? We did not build these models, the data scientists know them far better than we do, and we cannot read the weights.

All true. And none of it is new. So I ask back:

How do you challenge a credit model today? Have you ever rebuilt one to check it?

You have not, and you do not need to. You already challenge models no one in your team built. You do it through the first line's own validation, its limits, its reporting - laid out in a form you can interrogate. AI does not change the method. It raises the stakes on it, which is exactly why this is an extension of the model risk muscle you already have, not a new discipline.

You are not a technologist, and you are not a stamp

Two traps sit on either side of the second line, and firms fall into both.

The first is thinking you must out-build the data scientists to challenge them. Resist it. If your test for governing AI is whether your people can write a better model, you have set a test you will always fail, and the wrong one. Your object is not the model. It is the management of the model. Has this system been found and rated for how much it matters. Did someone independent of the builders validate it. Can a human actually intervene when it goes wrong. Is anyone watching it now it is live. Is someone in the business accountable for the outcome by name. Every one of those is a question you can press without reading a single weight.

The second trap is the opposite, and more common under deadline. You become the stamp. The business brings you a system the day before launch, you have no time and no evidence, and you sign because saying no is expensive. That is not the second line. That is a scapegoat with a signature. Real challenge means you can, and sometimes do, send a system back. If you never have, you are not the second line. You are a formality the first line routes around.

Evidence you can interrogate, not a demo

Here is where it goes wrong in the room. A first line wanting a fast yes will show you the system working. A smooth interface, a clean example, a confident data scientist. A demo is not evidence. It is the one path the business chose to walk you down, and it tells you almost nothing about the ninety-nine it did not.

What you are owed instead, and must demand, is evidence you can interrogate. The inventory entry. The materiality rating and the reasoning behind it. The validation from someone who did not build the thing. The test results, including on the cases where it failed. The monitoring plan for after go-live. And the definition of what "good enough" means for this task, with the test that proves the system clears it.

That last one is the quiet heart of it. Before any AI, the business should be able to tell you what a good answer and a bad answer look like for the job this system does. Fraud detection has to catch fraud without burying people in false alarms. A customer-facing tool has to fit the customer and not mislead. If the first line cannot say what good enough is for the task, no amount of model sophistication saves it, and no amount of your technical knowledge fills the gap. You are not assessing whether the AI is clever. You are making the business show that the job still gets done to standard, with the AI in the mix.

A demo is the one path the business chose to walk you down.

Why AI raises the stakes on an old method

The method is old. The reason it has to be applied harder is that an AI model has three properties ordinary software does not, and each one widens the gap between what the first line claims and what is true.

It is uncertain: it gives a spread of answers, not one, so "it works" is never the whole story, and you have to ask how often, and how badly, it does not. It is unexpected: nobody wrote its rules by hand, so it can find a route to the goal that no one designed, which is why your testing has to reach past the happy path into the edge cases. And it is opaque: the more capable the model, the harder it is to say why it did what it did, which is why reproducibility and documented validation matter more here than for a system whose logic you could just read. These are why your challenge probes behaviour, not logic. There is no rule to read, so you interrogate what the thing does.

None of that means you fear it. A model is a mathematical mapping from inputs to outputs. It does not understand, intend, or scheme. Treating it like a person, in either direction, is the fastest way to misread it - trust it like a colleague and you over-trust it, fear it like a villain and you go hunting for motives when the cause is a skewed training set. You hold it to account as what it is: a system the business chose to deploy, and must therefore be able to account for.

Good challenge is not a brake

One last thing, because the first line will tell you the opposite. Done properly, real challenge is not the thing that slows the business down. In the financial sector there is no return without risk, so managing risk is part of earning the return, not a tax on it. The business that has found its AI, knows what each system can do, can challenge the results and understand the thing, is the business that can move faster - deploy with confidence, say yes without a thousand nervous caveats. A solid base for evaluation and testing is not a blocker. A model nobody can vouch for is the blocker, because it either ships blind or stalls in doubt. When you challenge well, you are not holding the business back. You are part of how it gets to go.

For the second line

What to own. Your job is independent challenge and validation - not building the model, not rubber-stamping it. Judge the management of the model, not the maths: whether it is found, rated, validated by someone who did not build it, overseeable, monitored, and owned by name. Refuse the demo and demand evidence you can interrogate, above all a definition of "good enough" for the task and the test that proves it. Keep the authority to send a system back, and use it, or you are a formality. And remember challenge probes behaviour, not logic, because with AI there is no rule to read.

Ask the first line:

  • For this system, what does a good answer and a bad answer look like - and what test shows it clears the bar?
  • Show me the evidence, not the demo: the validation, the test results on the hard cases, the monitoring plan for after launch.
  • Who validated this who did not build it, and what could they have sent back?
  • If this system is wrong next week, who in the business answers for the outcome, by name?

Work with me

I train risk and compliance teams on turning AI risk management into a working system. See the courses and workshops, read more on AI risk management, or get in touch.