Study: Transparency is often lacking in datasets used to train large language models
The study reveals that the majority of large language model training datasets suffer from missing or incorrect licensing information, creating significant legal and ethical risks. When data origins and usage restrictions are lost during aggregation, practitioners may unknowingly use unsuitable or biased data, leading to unfair model predictions and potential compliance violations. This lack of transparency undermines the reliability of AI systems, particularly in sensitive areas like loan evaluations, where understanding data origins is crucial for assessing model limitations and risks. To address this, researchers created the Data Provenance Explorer, a tool that automatically generates clear summaries of dataset creators, sources, and licenses. This innovation directly supports open data initiatives by making opaque data ecosystems readable and navigable for the broader community. By clarifying who created the data and how it can be legally used, the tool empowers developers to make informed choices, ensuring that open datasets are utilized responsibly and effectively without requiring manual audits of thousands of sources. Furthermore, the audit highlighted a geographic bias in dataset creation, with most contributors located in the Global North, which limits model effectiveness in other regions. This finding emphasizes the need for more diverse and transparent data practices in the open data community. By exposing these disparities and providing tools to trace data heritage, the research encourages the development of more equitable AI systems. Ultimately, improving data provenance is essential for building trustworthy, legally sound, and culturally representative artificial intelligence models.
Source: news.mit.eduPublished on 2024-08-31
Related news
- Child abuse images removed from AI image-generator training source, researchers say
- Bolivia supera los 11 millones de habitantes y Santa Cruz es el departamento más poblado
- The org behind the dataset used to train Stable Diffusion claims it has removed CSAM | TechCrunch
- Regiones y políticos dudan de datos del Censo; Santa Cruz los rechaza