GPT-4o’s Chinese token-training data is polluted by spam and porn websites
The introduction of advanced tokenizers significantly impacts the efficiency and economics of multilingual large language models. By optimizing token lengths for non-English languages like Hindi, Bengali, Russian, and Arabic, systems can process prompts faster and reduce operational costs by nearly fourfold. This shift highlights a crucial implication for global accessibility, demonstrating that technical improvements can make high-quality AI services more affordable and viable for diverse linguistic communities, thereby expanding the practical reach of open AI technologies beyond English-centric development. However, this progress is unevenly distributed, revealing critical data quality disparities. While the model handles several low-resource languages effectively, Chinese tokens exhibit severe contamination, reflecting spam, pornography, and scams rather than legitimate linguistic structures. This inconsistency underscores a major vulnerability in training data curation, where insufficient cleaning efforts allow malicious content to permeate specific languages. Such flaws risk embedding harmful biases or inaccuracies into the model’s foundational knowledge, challenging the integrity of open models that rely on vast, uncurated internet scrapes. This article is vital to the open_data community as it exposes the hidden costs and ethical risks of relying on raw web data. It emphasizes that open data is not inherently clean or neutral; without rigorous transparency and cleaning standards, datasets can inherit and amplify societal harms. For developers and researchers advocating for open AI, this serves as a stark reminder that data provenance and quality control are not optional extras but essential prerequisites for building trustworthy, equitable, and safe public AI infrastructure.
Source: technologyreview.comPublished on 2024-07-08
Related news
- Google claims new AI training tech is 13 times faster and 10 times more power efficient — DeepMind's new JEST optimizes training data for impressive gains
- Projects funded under national AI programme AISG focus on good training datasets
- Revealed: Durham modules ranked by student attainment in 2021-2022
- Council paid out tens of thousands of pounds due to data breaches
- El I Foro Urbano de la Región conoce los 100 proyectos que conforman la Estrategia Murcia 2030 del Ayuntamiento