OpenAI has identified evidence that the Chinese startup DeepSeek employed a technique known as “distillation” to train its open-source model using outputs from OpenAI’s own models. This practice involves extracting knowledge from large, expensive models to create more efficient and cost-effective competitors. Although distillation is common in the industry, OpenAI considers it a violation of its terms of service, which explicitly prohibit using its outputs to develop products that directly compete with the company. The launch of DeepSeek’s models, which achieved performance comparable to leading U.S. models at a fraction of the hardware and energy costs, triggered a significant reaction in tech markets. This efficiency has raised concerns about the potential obsolescence of massive investments in high-performance computing infrastructure, temporarily affecting the valuation of key companies such as Nvidia. The situation highlights the growing ability of developers with limited resources to compete with tech giants through data optimization strategies. This conflict is central to the open-data ecosystem, as it illustrates the tense relationship between collaborative innovation and intellectual property protection. While some experts argue that leveraging the outputs of aligned models is a standard academic and commercial practice, OpenAI insists on the need to safeguard its technological assets. The case underscores the legal and ethical challenges of defining the boundaries of using derived data in AI training, impacting how companies can share or retain knowledge in a globalized market.
Source: df.clPublished on 2025-01-30