Quaintitative

MindForge Toolkit

Data management

Data management is making the data an AI uses fit for purpose, and keeping it governed through the system's life. It is one of the seventeen areas in the MindForge AI Risk Management Toolkit, the Singapore industry's practices for the AIRG. MindForge is the practices; the AIRG is the expectations and standards. I wrote the AIRG, and a model is only ever as good as the data it learned from, which is why this area carries so much of the risk.

What the AIRG expects

The AIRG expects data management controls that keep the data an AI uses fit for purpose and representative, of high quality, and under robust data governance (AIRG 4.5), with the depth proportionate to the system's materiality. For how data management sits alongside the other lifecycle controls, see the AIRG in practice, and the AIRG overview.

What MindForge says to do

The Toolkit sets out data management practices across the data an AI uses, both training data and the operational data a live system draws on, such as the content behind a retrieval-augmented generation (RAG) system.

Fit for purpose

  • Identify the sources and confirm they are relevant to and representative of the problem, collecting only what is needed.
  • Assess quality against the firm's existing data-quality dimensions, and apply those checks to the firm's own data even where a third-party model's training data cannot be inspected.
  • Assess representativeness and balance - a reasonable mix across attributes such as gender, race and language, and a range of tasks and edge cases the system will meet in operation.
  • Manage synthetic data against three tests: is it statistically representative, is it fit for purpose, and does it protect sensitive information.

Justify, document, control, and own

  • Justify personal attributes - identify each one used and document why, weighing performance against fairness and privacy, and record the privacy-preserving measures applied.
  • Document metadata and lineage for training and operational data, so a prediction can be traced back and the work can be validated and audited.
  • Control access to data, pipelines, configuration and model versions, with particular care for RAG, and guard against data poisoning with logging and change management.
  • Own the derived data - feature marts, vector indexes and other transformed data need a named owner accountable for quality and governance.
  • Identify and mitigate bias in training and test data, using bias-aware collection, disaggregated evaluation, tests such as disparate impact, and corrections such as re-weighting or re-sampling.

In practice

What good looks like. Data sources documented and representative of the population the system will meet, edge cases included; personal attributes justified or removed; lineage from source to model; access controls on data, pipelines and model versions; and bias tested and mitigated for the higher-risk use cases. The depth scales with materiality.

Evidence to hold:

  • A data-fitness and representativeness assessment, and the personal-attribute justification log.
  • Metadata and lineage documentation, and the data access-control matrix.
  • Bias-test results (such as disparate impact) with the mitigations applied, for material use cases.

How banks do it

In the MindForge Implementation Examples, a bank such as DBS applied data security and access controls, documented metadata and lineage through its data-onboarding process, and validated data quality against a golden set of questions and answers when deploying an internal generative AI assistant.

My take

Everything a model can do is a compressed record of the data it was trained on. Rubbish in, rubbish out is not a slogan here; it is the whole risk.

Firms spend heavily on the model and lightly on the data, which is backwards. The data decides what the model can and cannot do; the architecture mostly decides how fast. (From my AI Risk Management from First Principles primer.)

Work with me

I train and advise financial institutions on AI data management and the rest of the AIRG programme. See the courses and workshops, read more on AI risk management, or get in touch.