GPT-4o’s Chinese token-training data is polluted by spam and porn websites

The new tokenizer significantly reduces operational costs for non-English languages like Hindi and Bengali by enabling faster processing and shorter tokenization. This efficiency improvement benefits open data users by making multilingual large language models more accessible and affordable, allowing for broader inclusion of diverse linguistic datasets without prohibitive computational expenses. The quality of these tokens largely reflects standard news content, indicating that the underlying training data for these languages is relatively clean and useful. Conversely, the situation differs drastically for Chinese, where the tokenizer incorporates substantial amounts of spam, pornography, and scam-related content. Researchers highlight that the training corpus lacks sufficient cleaning, resulting in tokens heavily polluted by hijacked content and search engine manipulation tactics. This reveals critical vulnerabilities in how multilingual data is curated, showing that automated collection methods often fail to filter out malicious or irrelevant information effectively across all languages. This disparity is highly relevant to open data because it underscores the urgent need for rigorous data hygiene standards in public AI resources. If foundational models integrate corrupted data, the resulting systems perpetuate misinformation and bias, undermining the integrity of open-source knowledge. For developers and researchers relying on open data, this serves as a warning to prioritize high-quality, verified datasets and implement stricter cleaning protocols to ensure that open models remain trustworthy and reliable for global, multilingual applications.

Source: technologyreview.com
Published on 2024-05-18