Whether the data used to train an AI fairly reflects all the different types of people, situations, or events the AI will be used on — not just the easy or common cases.
Whether the data used to train an AI fairly reflects all the different types of people, situations, or events the AI will be used on — not just the easy or common cases.
Poor representativeness is a primary cause of algorithmic bias and performance degradation in deployment. A model trained on data that over-represents certain demographic groups, geographic regions, or time periods will perform unevenly — often failing for under-represented groups. ISO/IEC 5259 includes representativeness as a key data quality dimension for AI. The EU AI Act requires providers of high-risk AI systems to examine training data for representativeness and to document steps taken to address biases. Techniques to improve representativeness include stratified sampling, data augmentation, and targeted data collection.
AI teams building models for diverse user populations must explicitly measure and document representativeness — and engage domain experts to identify populations at risk of under-representation before data collection is finalised.
Like a clinical trial that recruits only young men — its results may not represent how the treatment works for women, elderly patients, or people of different ethnicities. Under-representative training data creates the same problem for AI systems.
Whether the data used to train an AI fairly reflects all the different types of people, situations, or events the AI will be used on — not just the easy or common cases.
Poor representativeness is a primary cause of algorithmic bias and performance degradation in deployment. A model trained on data that over-represents certain demographic groups, geographic regions, or time periods will perform unevenly — often failing for under-represented groups. ISO/IEC 5259 includes representativeness as a key data quality dimension for AI. The EU AI Act requires providers of high-risk AI systems to examine training data for representativeness and to document steps taken to address biases. Techniques to improve representativeness include stratified sampling, data augmentation, and targeted data collection.
AI teams building models for diverse user populations must explicitly measure and document representativeness — and engage domain experts to identify populations at risk of under-representation before data collection is finalised.