OpenAI’s Mira Murati is “not sure” where Sora’s training data comes from

OpenAI’s CTO Mira Murati admitted uncertainty regarding the precise data sources used to train the Sora video-generation model, stating only that it relies on publicly available or licensed information. This lack of transparency is significant because it highlights the ongoing ambiguity surrounding how major AI firms construct their foundational models. For open data advocates, this incident underscores the critical need for clear, accessible documentation of training datasets. The situation occurs amidst growing legal pressure and copyright lawsuits against OpenAI, including suits from authors and The New York Times. These challenges emphasize the tension between using vast internet data for AI development and respecting intellectual property and privacy rights. The reliance on undefined public data raises questions about accountability and the ethical sourcing of information. This ambiguity is particularly relevant to the open data community, which advocates for transparency in algorithmic processes. When data provenance is unclear, it becomes difficult to assess bias, legality, or ethical compliance. Ultimately, the Sora controversy illustrates why open standards for data usage are essential for building trust and ensuring responsible AI development in an increasingly regulated landscape.

Source: cointelegraph.com
Published on 2024-03-17