When a training dataset is lopsided in a way that causes an AI to produce unfair or inaccurate outputs for some groups or situations.
When a training dataset is lopsided in a way that causes an AI to produce unfair or inaccurate outputs for some groups or situations.
Dataset bias can arise from: historical bias (data reflecting past discriminatory patterns), representation bias (some groups collected more than others), measurement bias (data collection instruments that produce systematically different measurements for different groups), aggregation bias (grouping populations that should be modelled separately), and labelling bias (annotators applying inconsistent or discriminatory labels). Dataset bias is a root cause of algorithmic bias and is addressed in EU AI Act Article 10, which requires providers to examine training data for relevant biases and implement data governance measures to mitigate them.
AI ethics and compliance teams must build systematic bias auditing into data preparation processes — documenting identified biases and mitigation steps as part of Annex IV technical documentation.
Like a flawed medical study that recruited only patients from private hospitals — the resulting treatment recommendations will be subtly biased towards the health patterns of wealthier patients.
When a training dataset is lopsided in a way that causes an AI to produce unfair or inaccurate outputs for some groups or situations.
Dataset bias can arise from: historical bias (data reflecting past discriminatory patterns), representation bias (some groups collected more than others), measurement bias (data collection instruments that produce systematically different measurements for different groups), aggregation bias (grouping populations that should be modelled separately), and labelling bias (annotators applying inconsistent or discriminatory labels). Dataset bias is a root cause of algorithmic bias and is addressed in EU AI Act Article 10, which requires providers to examine training data for relevant biases and implement data governance measures to mitigate them.
AI ethics and compliance teams must build systematic bias auditing into data preparation processes — documenting identified biases and mitigation steps as part of Annex IV technical documentation.