GPT-4o’s Chinese token-training data is polluted by spam and porn websites

The introduction of advanced tokenizers significantly impacts the efficiency and economics of multilingual large language models. By optimizing token lengths for non-English languages like Hindi, Bengali, Russian, and Arabic, systems can process prompts faster and reduce operational costs by nearly fourfold. This shift highlights a crucial implication for global accessibility, demonstrating that technical improvements can make high-quality AI services more affordable and viable for diverse linguistic communities, thereby expanding the practical reach of open AI technologies beyond English-centric development. However, this progress is unevenly distributed, revealing critical data quality disparities. While the model handles several low-resource languages effectively, Chinese tokens exhibit severe contamination, reflecting spam, pornography, and scams rather than legitimate linguistic structures. This inconsistency underscores a major vulnerability in training data curation, where insufficient cleaning efforts allow malicious content to permeate specific languages. Such flaws risk embedding harmful biases or inaccuracies into the model’s foundational knowledge, challenging the integrity of open models that rely on vast, uncurated internet scrapes. This article is vital to the open_data community as it exposes the hidden costs and ethical risks of relying on raw web data. It emphasizes that open data is not inherently clean or neutral; without rigorous transparency and cleaning standards, datasets can inherit and amplify societal harms. For developers and researchers advocating for open AI, this serves as a stark reminder that data provenance and quality control are not optional extras but essential prerequisites for building trustworthy, equitable, and safe public AI infrastructure.

Source: technologyreview.com
Published on 2024-07-08