Tech Thoughts: Media needs a united front against data scraping to train AI
The article argues that the current trajectory of artificial intelligence development disproportionately benefits a handful of large technology firms by harvesting content from millions of users without adequate consent. While regulatory frameworks in some regions propose "opt-out" mechanisms for data scraping, this approach places the burden on creators to protect their work, rather than requiring explicit permission. Consequently, most individuals and smaller entities remain exposed to having their digital footprints used to train proprietary AI models, exacerbating existing power imbalances in the digital economy. A critical failure in the media sector’s response is the fragmentation of alliances, exemplified by The New York Times withdrawing from a broader coalition of news organizations. By choosing to negotiate individually, major outlets bypass opportunities to unite with smaller publishers who are often more vulnerable to exploitation. This splintered approach prevents the creation of a unified front capable of demanding fair compensation or strict licensing agreements, leaving the industry divided and significantly weaker in its stance against unauthorized data mining. This lack of solidarity has profound implications for open data and digital rights, as it entrenches a system where only well-resourced entities can dictate terms to tech giants. If media groups fail to collaborate, AI companies will likely continue to prioritize and fund those with the largest datasets, effectively monetizing public information while sidelining independent journalism. True fairness requires a collective push for opt-in standards or universal compensation, ensuring that the value generated by AI reflects the diverse contributions of the entire information ecosystem rather than just a few privileged partners.
Source: rappler.comPublished on 2023-08-17