ChatGPT 5 está listo: de dónde obtuvo los datos públicos

OpenAI’s introduction of GPTBot represents a significant evolution in data collection for artificial intelligence, enabling the massive indexing of public information to train its next-generation models. This tool operates under a voluntary opt-out system, where website owners must explicitly configure rules to prevent their sites from being crawled, contrasting with passive consent methods. From an open data perspective, this approach raises ethical concerns about privacy and informed consent, particularly given the company’s track record. While OpenAI prioritizes controlled data accumulation to enhance its commercial products, competitors like Meta promote a more transparent open-source model, albeit with restrictive licenses. This difference reflects two opposing philosophies: one focused on closed systems and monetization, versus another that fosters community access and customization. The relevance of this article lies in how these changes in data collection affect accessibility and transparency within the AI ecosystem. The ability to scan and process vast amounts of public information without active consent challenges fundamental principles of data governance. This directly impacts the creation of open digital resources, setting precedents regarding who controls and benefits from data extracted from the web—a crucial issue for the future of open technology.

Source: infobae.com
Published on 2023-09-02