OpenAI wants to work with organizations to build new AI training data sets | TechCrunch

The article highlights the critical issue of bias and toxicity in AI training data, which often reflects Western-centric viewpoints and harmful content. OpenAI has launched a new initiative to collaborate with external institutions to curate broader, more representative datasets. This effort aims to expand AI’s understanding of diverse cultures, languages, and specialized domains, addressing the flawed foundations that currently limit model utility and safety. This partnership model seeks to digitize and aggregate large-scale data that is not easily accessible online, creating both open-source and private datasets. By including diverse content, OpenAI intends to make its models more helpful and universally applicable. The approach involves working with governments and organizations to improve language capabilities and domain-specific knowledge, suggesting a shift towards more inclusive and comprehensive data collection methods for future AI development. This initiative is highly relevant to open data because it underscores the urgent need for diverse, high-quality, and ethically sourced information to train fair AI systems. However, the article also raises concerns about commercial motives and lack of compensation for data contributors, echoing ongoing debates about intellectual property and consent. It serves as a reminder that while expanding data access is crucial, ensuring equitable practices and transparency in data partnerships remains a significant challenge in the open data landscape.

Source: techcrunch.com
Published on 2023-11-10