Harvard and Google to release 1 million public-domain books as AI training dataset | TechCrunch

Harvard University is launching the Institutional Data Initiative to democratize access to vast public-domain book datasets for AI training. Backed by Microsoft and OpenAI, this project aims to lower barriers for researchers and startups, reducing the financial advantage currently held by large tech firms. This effort directly supports open data by creating a trusted, legal conduit for public-domain texts. It ensures that critical training resources are accessible to the broader scientific community rather than being monopolized by deep-pocketed corporations, fostering a more equitable AI development landscape. By making millions of books available for machine learning, the initiative promotes transparency and inclusivity in artificial intelligence research. This alignment with open data principles encourages wider participation and innovation, ensuring that the benefits of advanced AI are not restricted to a few well-funded entities.

Source: techcrunch.com
Published on 2024-12-13