Quaintitative

AI Supervision

Assessing AI controls before deployment

Before an AI system goes live, the firm builds it. A supervisor's job is not to build alongside them. It is to judge whether what they built was built to a standard you can see, and whether anyone independent checked before the switch was thrown. I supervised AI at the Monetary Authority of Singapore and wrote the AIRG, and when it comes to assessing AI controls before deployment, this is the order I work in. Faced with the list of development controls, data, selection, explainability, fairness, testing, security, documentation, validation, supervisors do one of two things. They tour all of them shallowly, or they fixate on one and miss the rest.

There are so many lifecycle controls. Do we really have to assess every one of them for every system?

No, and trying would be the proportionality mistake. So the question sharpens: which of these controls, if it is weak, makes all the others unreliable, and which one, done well, tells you most about the whole build? The controls are not a flat checklist to tick through. They chain. Bad data makes a good model unsound; an untested model makes every claim about it a guess; and nothing a firm did before go-live counts for much if no one independent of the builders checked it. What follows walks the development controls in that spirit, what each is for and where it quietly fails, with the evaluation and the independent gate as the two that carry the most weight.

Data: fit for purpose, and representative

A model is its training data wearing a function. If the data was wrong, the model is wrong in ways no later control fully fixes, so this is where you look first in the build.

The AIRG wants data that is fit for purpose, representative, of good quality, and under real governance. Representativeness is the one to press: does the training data actually match the population and the conditions the system will meet in production, or was it whatever happened to be on hand. A credit model trained on a narrow slice of borrowers will fail quietly on everyone outside it. Then quality, completeness, accuracy, timeliness, with the plain red flags: heavy missing data in the features that matter, labels of unknown reliability, external data taken on trust from a vendor without due diligence. And lineage: can the firm trace a given output back through the transformations to the source, which is what makes an audit or an incident investigation possible at all. You are not auditing their data science. You are checking whether they can show the data was fit, representative, and traceable, or whether "the data is fine" is an assertion with nothing behind it.

Selection: justify the complexity

The AIRG asks developers to justify and document why they chose the model they did, especially when they reach for a more complex one. This is a control supervisors under-use, and it is a good one.

The question to put is whether a simple baseline was built and beaten. The honest process develops a plain model first, measures it, and only moves to something more complex if the gain is real and worth the cost in opacity and risk. Where there is no baseline, there is usually no discipline, just a reach for the most powerful thing available, with the complexity accepted by default rather than chosen. A marginal accuracy gain bought with a large loss of explainability is a trade the firm should have made consciously, and should be able to defend to you. If they cannot tell you what the simpler option would have cost them, they did not really choose.

Transparency and explainability: enough to act on

The AIRG ties transparency to materiality, more for high-stakes decisions, lighter for low-risk uses, and judges it by whether the explanation lets someone make an informed decision, not by whether a technique was applied.

So two things. For customers and affected people: is the disclosure explicit that AI is being used, in plain words, and does an adverse decision come with a reason the person could act on, what would have to change for a different outcome, rather than "computer says no" dressed in legal language. For the firm's own people: do the humans meant to oversee the system get explanations good enough to actually challenge it. A SHAP chart no reviewer understands is not explainability; it is decoration that passes an audit. Match the depth to the audience and the stakes, and test it by asking to see a real explanation that was given to a real customer or reviewer.

Fairness: defined, tested, and traded off in the open

The AIRG asks the firm to define what it considers fair for its context, and to have controls to find and mitigate harmful bias, proportionate to materiality. The trap here is a fairness policy that never touches a model.

Check that the firm has actually chosen a fairness definition and metric suited to the use case, and can say why, because the different definitions genuinely conflict and you cannot satisfy them all at once. Check that it has identified not just the obvious protected attributes but the proxies that stand in for them, since a model can discriminate through a postcode without ever seeing a protected characteristic. Check that bias was tested before deployment, across sub-populations, with sample sizes large enough to mean something. And check that where fairness cost some accuracy, that trade was made in the open and signed off, not buried. For a lending or pricing model, this is not a soft control, it is where the firm's regulatory and conduct exposure concentrates.

Evaluation is the one control that can move all the others.

Evaluation and testing: the control that moves the rest

If you press one development control hard, press this one, because it is upstream of almost everything else you care about. The AIRG wants evaluation and testing proportionate to materiality, across a range of plausible conditions from the ordinary to the edge cases. Done well, it is what tells the firm, and you, how uncertain the system is, where it breaks, and whether it clears the bar for the task. Done badly or not at all, every other claim about the system is a guess.

The failures are specific and worth knowing. Test data that leaked from training, so the reported performance is a flattering illusion. Testing only on the happy path, with no edge cases, no stress, no underrepresented groups. For systems where it applies, no adversarial or red-team testing, no probing for prompt injection or jailbreaks in a generative system, no evasion testing in a classifier. And, underneath it all, no stated threshold: no definition of what "good enough" means for this task, so the test cannot pass or fail, only produce numbers. That last one is the quiet heart of the matter, the same point as the whole book: if the firm cannot tell you what a good answer and a bad one look like for the task, the testing proves nothing, however much of it there is. Ask for the test results on the cases where the system did badly. A firm that only has results where it did well has not finished testing.

Security, and the record that makes it all checkable

Two controls here do quiet, load-bearing work. The first is technology and cyber: an AI system is still a system, and it has to be secure and well-governed against the ordinary technology risks plus the AI-specific ones like data poisoning. The second is reproducibility and auditability, which is the control that makes your job possible at all. The AIRG wants development documented well enough that a competent, independent reviewer could rebuild the work, the data versions, the code, the training procedure, the configuration, archived so a production model can be traced to the exact run that produced it. This is not bureaucracy. Without it, "we validated it" cannot be verified, an incident cannot be investigated, and a model that drifted cannot be compared to what it used to be. A firm that cannot reproduce its own model has nothing for you, or its own second line, to actually check.

The gate: someone independent, before go-live

Everything above is the build. This is the gate, and it is where before-go-live supervision concentrates. The AIRG requires that an AI use case be reviewed before deployment by parties not involved in building it, and that high-materiality systems undergo formal independent validation. Not a peer glancing over a colleague's work. Independence with teeth.

So you check three things about the gate. That the reviewer was genuinely independent of the builders, reporting through a different line, with no stake in the system shipping. That the review had the authority and the depth to matter, that it could, and sometimes did, send a system back or attach conditions, proportionate to the risk. And that where approval came with conditions, those conditions were actually tracked and closed before or shortly after go-live, rather than noted and forgotten. The most revealing request here: show me a system your validation function sent back, and what it demanded. A validation function that has approved everything it has ever seen is not a gate. It is a turnstile, and the firm walks through it at will.

For the supervisor

What to look for. The development controls chain, so assess them as a chain, proportionate to materiality. Data that is representative of production and traceable to source, not just "fine." A documented reason the firm chose the model it did, measured against a simple baseline it actually built. Transparency deep enough for a customer to act on and a reviewer to challenge, not a technique applied for show. Fairness with a chosen definition, proxy variables considered, pre-deployment testing across sub-groups, and any accuracy trade-off made in the open. Above all, evaluation and testing with a stated "good enough" threshold, no train-test leakage, real edge-case and adversarial testing, and results that include the failures. Security against AI-specific attacks, and documentation good enough to reproduce the model. And the gate: genuinely independent pre-deployment validation for material systems, with the authority to send a system back and a record that it sometimes has.

Ask the firm:

  • How did you confirm the training data represents the population this system meets in production, and can you trace an output back to its source data?
  • What simpler model did you build as a baseline, and what did the more complex one buy you over it?
  • Show me a real explanation this system gave a customer or a reviewer. Could they act on it?
  • For a lending or pricing model: what fairness definition did you choose and why, which proxy variables did you test, and what did fairness cost you in accuracy?
  • What does "good enough" mean for this task, what threshold did you set, and show me the test results on the cases where the system did worst.
  • Can you reproduce this model from your records, the data version, code, and configuration?
  • Show me a system your independent validation function sent back or made conditional, and whether those conditions were closed.

Work with me

I train regulators, supervisors, and public authorities on AI governance and risk management. See the courses and workshops, read more on AI risk management, or get in touch.