Research shows AI image generators could be their own demise
Recent research reveals that training AI image generators on synthetic data leads to a rapid deterioration in output quality, a phenomenon akin to genetic inbreeding. As models consume their own generated images, the resulting "AI cannibalization" causes collapse, rendering future iterations increasingly nonsensical. This inherent limitation suggests that without human-created source material, generative AI cannot sustainably improve or adapt to new trends. The prevalence of AI-generated content online exacerbates this risk, as scrapers inevitably harvest these images for training. Consequently, developers may face rising costs to identify and filter out synthetic data, or they might discover that current human-centric datasets are already sufficient for maintaining quality. This highlights a critical dependency on authentic human creativity, as models cannot evolve indefinitely by feeding on their own outputs. This article is relevant to open data because it underscores the necessity of data provenance and curation. It implies that open datasets used for AI training must be rigorously filtered to exclude synthetic content to prevent model degradation. Furthermore, it supports the development of open standards for content credentials, enabling the community to distinguish between human and machine-generated work, thereby preserving the integrity of open data resources for future technological advancement.
Source: creativebloq.comPublished on 2024-08-06