Using bigger AI training data sets may produce more racist results

Contrary to the widespread belief that scaling up data improves fairness, recent research indicates that larger training sets can actually exacerbate racial biases in artificial intelligence. By comparing datasets of vastly different sizes, researchers found that models trained on more extensive data were significantly more likely to associate Black individuals with criminal categories. This challenges the industry assumption that quantity inherently leads to diversity, revealing that simply adding more internet-scraped data often amplifies existing prejudices rather than mitigating them. The study highlights that larger datasets frequently contain higher proportions of hateful or aggressive content, suggesting that current scaling practices are counterproductive to ethical AI development. While some organizations argue the findings are overstated, the core implication remains that unchecked data expansion without rigorous quality control leads to worse outcomes. This exposes a critical gap in the field, as many tech companies fail to perform basic checks to remove biased samples, treating scale as a sufficient solution to complex social issues. This article is vital to open data advocates because it underscores the urgent need for transparency and accountability in AI training resources. It critiques major corporations for using closed, unverifiable datasets that may be even more biased than open-source alternatives, while simultaneously calling on open-data providers to improve their quality standards. Ultimately, it argues that the open data community must prioritize rigorous auditing and public scrutiny over mere volume to ensure AI systems are equitable and trustworthy.

Source: newscientist.com
Published on 2023-12-05