Quaintitative

MindForge Toolkit

Data ethics

Data ethics is the question of whether the data a firm feeds its AI is lawful, consented, and fair to use in the first place, before anything is built with it. It is one of the seventeen areas in the MindForge AI Risk Management Toolkit, the Singapore industry's practices for the AIRG. MindForge is the practices; the AIRG is the expectations and standards. I wrote the AIRG, and this guide sets out the practice, the expectation it meets, and the evidence it produces.

What the AIRG expects

The AIRG expects the data used in an AI use case to be fit for purpose and representative, and harmful bias to be identified and mitigated, in proportion to the system's materiality (AIRG 4.5 on data management and 4.8 on fairness). The ethics of the data, whether it is lawful and fair to use at all, sits underneath both. For how this fits with the rest of the controls, see the AIRG in practice, and the AIRG overview.

What MindForge says to do

The Toolkit frames data ethics as a check, done before the data is used, that its intended use is compatible with ethical, regulatory and organisational standards.

Lawful, consented, and fair to use

Identify the data obligations that apply, given where the firm operates, where the data subjects are, how the data is collected and processed, any cross-border transfers, and the contracts involved. Apply existing data policies to all data the use case touches. Sensitive or high-impact data, such as health data, biometrics, or private customer and employee information, calls for extra controls and a supplementary ethics review and escalation. Where customer data is used to train AI, obtain specific consent for that use, and include privacy-preserving measures.

Third-party data on the right terms

Data brought in from a third party has to respect intellectual property rules, contractual obligations and licensing rights, and should be used only for the purposes it was obtained for. Where a generative AI system was trained on externally sourced creative content, confirm that use complied with the rules in the jurisdictions the firm operates in.

In practice

What good looks like. Before a use case is built, someone has checked the data is lawful to use - consent for customer data, licensing for third-party data, the cross-border and contractual position - and sensitive datasets have been through an ethics review and escalation, proportionate to the risk. The check happens before the build, not as a sign-off after it.

Evidence to hold:

  • A data-obligations assessment for the use case, covering jurisdictions, data subjects, cross-border transfers, and contracts.
  • Consent records and privacy-preserving measures where customer data is used for training.
  • Licensing and IP checks for third-party and training data, and ethics-review records for sensitive datasets.

How banks do it

The MindForge Implementation Examples show the pattern: a bank such as DBS runs a responsible-data-use framework, where a cross-functional responsible-AI group and a data council review a use case's data before it ships, with location-specific legal and compliance sign-off when the system rolls out across markets.

My take

A model learns the world you show it, biases included. Decide what data you are willing to use, and whether it looks like the world the model will act in, before you build, not after.

Data ethics is cheap to do early and expensive to retrofit. The consent you did not get and the licence you did not check do not surface until the system is live and in front of a customer. (From my AI Risk Management from First Principles primer.)

Work with me

I train and advise financial institutions on data ethics and the rest of the AIRG programme. See the courses and workshops, read more on AI risk management, or get in touch.