Using bigger AI training data sets may produce more racist results

Contrary to the prevailing industry belief that scaling data volumes inherently reduces prejudice, recent research demonstrates that larger training sets can actually exacerbate AI bias. By comparing datasets of varying sizes, researchers found that models trained on significantly more data were substantially more likely to associate Black individuals with criminal categories. This challenges the assumption that quantity leads to quality, revealing instead that the sheer volume of internet-scraped data often amplifies existing societal prejudices rather than diluting them. This phenomenon occurs because increased scale does not guarantee diversity; instead, it frequently incorporates a higher concentration of hateful or aggressive content found on specific subsets of websites. The study indicates that larger datasets statistically contain more targeted speech, leading to more discriminatory outcomes in facial classification tasks. Consequently, simply adding more data without rigorous curation fails to address fundamental biases, suggesting that the current "bigger is better" paradigm may be actively harming the fairness of open-source and commercial AI systems. The findings are critical for the open_data community, as they highlight the urgent need for basic quality checks and transparency in training materials. While open datasets allow for scrutiny, many corporations rely on closed, unexamined data that may be even more biased. This underscores the importance of advocating for open, auditable data practices, as public visibility is the only current mechanism to identify and mitigate these dangerous scalability pitfalls in artificial intelligence.

Source: newscientist.com
Published on 2023-07-20