Terms for managing, measuring, documenting, and improving data used in AI systems.
A searchable directory of all the data an organisation has — describing what each dataset contains, where it lives, how good it is, and who owns it.
A missing or thin patch in a dataset — where certain types of people, places, or situations are not captured or not captured enough.
The rules and responsibilities that decide who can do what with data inside an organisation, and how that data should be kept accurate and secure.
The complete journey of data — from when it is first created or collected, through how it is used and stored, to when it is eventually deleted.
A record that shows where data came from, what happened to it along the way, and where it ended up — like a passport stamp history for data.
The day-to-day activities involved in handling data — collecting it, storing it safely, keeping it accurate, and retiring it when no longer needed.
Collect only the data you actually need for the job, and don't keep it longer than necessary.
All the work done to clean and organise raw data before it is used to train an AI — removing errors, standardising formats, and making sure it is in a usable shape.
Documentation of where a piece of data originally came from, who owned it, and who has handled it — proving it is genuine and untampered.
How good data is for the job it needs to do — whether it is accurate, complete, up to date, and not misleading.
One of the key aspects of data quality — like accuracy (is it correct?) or completeness (is it all there?) — used to frame what makes data good or bad for a given purpose.
A formal system of processes and controls that an organisation puts in place to set, measure, and maintain data quality standards.
A specific way of measuring how good data is — for example, the percentage of records with a complete address field, or the proportion of duplicate entries.
Whether the data used to train an AI fairly reflects all the different types of people, situations, or events the AI will be used on — not just the easy or common cases.
A person responsible for keeping particular data assets accurate, well-documented, and properly used within their part of the organisation.
Taking responsible, ongoing care of data assets — keeping them accurate, well-described, and handled in line with the organisation's rules and values.
Checking data to make sure it follows the right format and meets expected rules before it is used — like confirming that a date field actually contains a valid date.
Checking that data is not just correctly formatted but actually correct — confirming the values match what is true in the real world.
When a training dataset is lopsided in a way that causes an AI to produce unfair or inaccurate outputs for some groups or situations.
Data is 'fit for purpose' when it is good enough for the specific job you need it to do — quality is judged against the task, not in the abstract.
Keeping one agreed, accurate version of important business records — like a single master customer list — so that every system in the organisation uses the same information.
Data about data — labels and descriptions that tell you what a dataset contains, where it came from, how reliable it is, and how to use it correctly.
The practice of keeping all the descriptions and context about data accurate, organised, and up to date across an organisation.