Can AI even be open source? It's complicated

The article argues that while the fundamental software frameworks driving artificial intelligence are open source, the most popular large language models are not. Despite marketing claims from major tech companies, these models often withhold critical components like training data and weights, a practice the author terms "open-washing." This discrepancy creates a misleading impression of transparency and collaboration, as the core intelligence remains proprietary and inaccessible for true inspection or modification. This situation is relevant to open data because it highlights the insufficiency of traditional software definitions when applied to AI. Neural network weights and extensive training datasets do not fit neatly into existing open-source licenses, which were designed for human-readable code. Consequently, there is an urgent need to redefine what constitutes "openness" in the context of machine learning artifacts to ensure genuine transparency. The resolution involves debating whether access to model weights alone is sufficient or if full access to training data is required for true openness. Advocates stress that without data, users can only tweak models rather than understand or reproduce them, undermining the open-source principle of deep inspection. A new, workable definition is being developed to balance commercial interests with the community’s need for verifiable, transparent AI development.

Source: zdnet.com
Published on 2024-08-06