Tumblr's Parent Company Enters Agreements with OpenAI, Midjourney for Training Data
Automattic, the parent company of Tumblr and WordPress, is finalizing agreements with major AI firms like OpenAI and Midjourney to license user-generated content for training models. While this reflects a growing industry trend of monetizing online data, Automattic plans to introduce opt-out mechanisms for users, signaling an attempt to balance commercial interests with user autonomy amidst significant privacy concerns. A critical complication involves an inadvertent initial collection of vast amounts of historical Tumblr data, raising serious questions about whether non-public or opted-out content was inadvertently shared. This incident highlights the risks inherent in scraping practices and underscores the urgent need for transparent data governance. It serves as a cautionary tale for the open data ecosystem, demonstrating how poorly managed data handling can erode trust and potentially violate the principles of open access and user consent. This development is highly relevant to open data because it illustrates the tension between the availability of public digital content and the ethical responsibilities of data stewards. As companies increasingly commercialize open web data, it becomes crucial to establish clear standards for attribution, consent, and exclusion. The situation emphasizes that open data initiatives must evolve to protect creator rights, ensuring that the availability of data does not come at the cost of individual privacy or the exploitation of creative communities.
Source: techtimes.comPublished on 2024-02-29
Related news
- Reconocen los derechos de autor frente a los servicios de IA
- Estados Unidos limitó el acceso de China y Rusia a los datos sensibles de los estadounidenses
- Law Enforcement Center Authority responds to KTIV’s story on our Freedom of Information request
- How synthetic data powers AI innovation – and creates new risks