We should all be worried about AI infiltrating crowdsourced work

Researchers from EPFL have uncovered that a significant portion of crowd workers on Amazon Mechanical Turk use generative AI tools like ChatGPT to complete tasks. This practice undermines the core value proposition of these platforms, which rely on genuine human insight for tasks that algorithms currently cannot handle effectively. The widespread adoption of AI by workers suggests that a substantial fraction of collected data may not reflect the intended human gold standard, but rather machine-generated content. This discrepancy poses a critical threat to the development of reliable artificial intelligence models. Data scientists typically treat human-generated data differently from LLM-produced data, as training models on synthetic text can amplify biases and perpetuate errors. If datasets intended for benchmarking or training contain undisclosed AI contributions, the resulting models may inherit these flaws, leading to unreliable outputs and invalid performance metrics. The difficulty in distinguishing between human and machine text exacerbates this issue, making it hard to ensure data integrity. This finding is highly relevant to the open_data community because it highlights a fundamental challenge in maintaining data provenance. Open datasets often serve as the foundation for transparent and reproducible AI research, yet they may inadvertently include "contaminated" data from undetected AI usage. Ensuring the authenticity of contributors is essential for preserving the utility of crowdsourced information. Without rigorous methods to verify data origins, the value of open datasets in advancing trustworthy AI technologies is significantly compromised.

Source: techcrunch.com
Published on 2023-06-19