GPT-4o’s Chinese token-training data is polluted by spam and porn websites
The introduction of a new tokenizer significantly reduces operational costs for large language models in non-English languages like Hindi and Bengali by enabling faster processing of longer, contextually relevant tokens. This efficiency allows providers to charge less for identical outputs without necessarily dramatically improving linguistic quality, highlighting how technical infrastructure improvements can directly impact accessibility and economic barriers for multilingual users. Conversely, the data reveals severe quality disparities, particularly with Chinese tokens, which are heavily polluted by spam, pornography, and scam-related content. Researchers note that the training corpus lacks adequate cleaning efforts for these specific languages, suggesting that content farms and search engine manipulations have contaminated the foundational data. This contrast indicates that while some languages benefit from rudimentary but clean web structures, others suffer from aggressive digital pollution that was not properly filtered before model training. This article is highly relevant to open_data because it underscores the critical importance of data hygiene and transparency in AI development. It demonstrates how uncurated, polluted datasets can permanently embed toxic content into foundational models, making post-hoc corrections difficult. For the open data community, this serves as a warning that simply accessing or utilizing large-scale web data without rigorous, language-specific cleaning protocols risks perpetuating harmful biases and security vulnerabilities in public AI tools.
Source: technologyreview.comPublished on 2024-06-01
Related news
- Hugging Face says it detected 'unauthorized access' to its AI model hosting platform | TechCrunch
- Freedom of Information Act
- Clearing Rights For A ‘Non-Infringing’ Collection Of AI Training Media Is Hard
- Inflación en Estados Unidos permanece estable en abril
- Un centro estadístico independiente que concentre todos los datos públicos ayudaría la democracia