Quaintitative

Podcast

Making AI safe in practice

On the AI Trust Talks podcast, I was asked to cut through the hype and talk about what actually makes AI safe in practice. I wrote the AIRG while leading AI risk supervision at the Monetary Authority of Singapore, and before that I spent years on model risk. My short answer: safety is not a thing you add at the end. It is the controls themselves, run as a system, and most of the fundamentals are ones the financial sector has used for decades. Here are the main points, and you can watch the full conversation below.

Safety is not a bolt-on

The biggest misconception is that you build the AI first and then bolt safety on afterwards. Safety is the controls: evaluation and testing, explainability where it is needed, human oversight, monitoring. Skip them and you end up in proof-of-concept hell, where everything demos nicely but no one dares sign off, because nobody has tested it and nobody has any real confidence in it. The controls are not the brake on getting something live. They are what lets you get it live.

The model-risk lens, and what the new AI widened

An AI system is a software system. What makes it different is the model. So the questions are the ones model risk has always asked: does it work as intended, which is what evaluation and testing are for; is it being used where it should be, because a model tested for one use case should not be dropped into another; and are the controls proportionate, because you do not apply every control uniformly, and sometimes you do not need explainability at all. Generative and agentic AI did not change those questions. They widened the envelope - more uncertainty, more unexpected behaviour, more opacity - which makes the same job harder, not different.

That is also why I would not treat a model's "reasoning" as the gold standard of explainability. Whether it comes from chain-of-thought prompting or a model trained to produce it, it is still the same step-by-step generation underneath, and it can still be a fabrication. For agentic systems the more useful thing is that you can decompose the steps and log what the system did at each one. That log is an audit trail you can actually examine, even when the step itself cannot be fully explained.

Benchmarks are someone else's test. Validation is yours.

Benchmarks are a public test, and these days they are often folded into the training data, so a model looks good on them by construction. They cannot tell you whether a model works in your context. Validation can, because it is your own test, on your own use case, inside your own organisation. Build those in-context test sets over time and they become a real advantage: when a new model lands - and they land weekly now - you can run it through your sets and know quickly whether it works for you. Taken further, those same test sets become fine-tuning data tuned to your context. The thing people treat as a compliance chore turns into how you adopt new models faster than anyone relying on benchmarks.

The controls are interlocked, not a checklist

Most people treat the controls as standalone boxes: explainability, fairness, monitoring, each ticked on its own. Working on the guidelines convinced me they are interlocked. You cannot monitor well until you have evaluation and testing, because testing is what tells you what the monitoring thresholds should be. You do not want explainability for its own sake; you want it so a human can oversee the system. See them as linked and you can solve problems you could not solve in isolation. A model is not explainable, which will often be the case, but you can still test its behaviour across representative cases and design human oversight to cover the gap, and that gets you somewhere you are comfortable deploying from.

The real blind spot: the AI you cannot see

Ask where the risk hides and it is not hallucination or prompt injection first. It is being blind to the AI you actually use. Firms that cannot track where AI is used, and how, get stuck in a loop of rebuilding the inventory. And recording it does not mean a name and a use case; it means what it is used for, how it was tested, and its known weaknesses. That was the fundamental failure in traditional model risk, and it is harder now, because AI is no longer confined to credit, market, and liquidity models - it is everywhere. Get that first step wrong and nothing downstream is trustworthy.

Two failure modes worry me more than the headline ones. First, stopping at principles: a responsible-AI statement on the website and nothing where the rubber meets the road, which looks like governance and is the opposite of it. Second, skill atrophy - the more we lean on generative AI, the more the judgement needed to sense when something is off drains away, and that judgement is exactly what using these tools well requires.

Make it a flywheel, not a finish line

The boring, underrated risk is seeing a demo work once and assuming it works everywhere. A thing that works in a single demonstration is the most dangerous kind of "done", because it ships without testing and without monitoring after go-live. The recent agentic-AI security incidents make the point: the AI did not cause them - exposed access keys and over-privileged credentials did. People playing with the shiny new toy forgot the old fundamentals still mattered.

So operationalising safety is not a certificate you earn once and file. It is a flywheel: identify and rate your AI, test it, set monitoring from the test, and when it drifts past the thresholds, run the loop again. Balance innovation against safety through proportionality - decide per use case what controls it actually needs to sit inside your risk appetite - and remember the incentives are aligned: the developer wants the thing to work, and so does the risk manager. Treat it as a conversation between them, not one side building and the other side checking. And do not reinvent the wheel: build on your enterprise risk management, because the fundamentals already cover most of it.

Work with me

Watch the full conversation on the AI Trust Talks podcast. I train and advise financial institutions on making AI safe in practice, grounded in AI risk management, the AIRG, and governing agentic AI. See the courses and workshops, or get in touch.