Las webs secretas con las que se entrena ChatGPT para parecer 'inteligente'
The popularity of artificial intelligence chatbots has skyrocketed, fueled by the massive ingestion of text scraped from the internet. This analysis reveals that their training datasets are dominated by sites focused on journalism, entertainment, and software development, with platforms such as Wikipedia and patent repositories serving as foundational pillars. This reliance on public sources defines the model’s linguistic imitation capabilities, although it lacks genuine understanding. The composition of the data raises critical implications regarding the quality and neutrality of the generated information. The predominant inclusion of news portals and personal blogs can propagate biases or misinformation, making it difficult to trace original sources. Moreover, the unbalanced representation in religious categories and the exploitation of ideas on crowdfunding platforms raise questions about the ethics and fairness inherent in automated training processes. This research is essential for the open data community, as it demystifies transparency in the lifecycle of AI models. By exposing what proprietary or personal information is used without explicit consent, it underscores the urgency of regulating web data collection. It thus fosters debate on integrity, privacy, and responsibility in accessing the public information required to train technologies that profoundly impact society.
Source: 20minutos.esPublished on 2023-05-30