Why Most AI Bias Benchmarks Measure the Wrong Thing
Researchers and developers often misreport AI bias by conflating four distinct types: representational harm, allocative harm, performance disparity, and viewpoint slant, each requiring different measurements and remedies. Common benchmarks like WEAT, StereoSet, and CrowS-Pairs have been criticized for flawed item construction, meaning scores may not reflect the behaviors they claim to measure. A landmark review by Blodgett and colleagues identified significant problems with widely used stereotype benchmarks, including non-minimal pairs and contested attributions. Results from these tools also shift depending on prompt templates, option ordering, and decoding settings, making cross-model comparisons unreliable. Experts argue that only downstream audits of real pipelines, using paired inputs that vary solely by group signal, can meaningfully detect allocative harm relevant to real-world deployment.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in