Google joins the war on AI hallucination with its massive Data Commons knowledge graph
Large language models frequently produce hallucinated, incorrect information because they function as predictive tools rather than reasoning entities. This persistent inaccuracy hinders the development of truly reliable AI. To address this challenge, Google has introduced DataGemma, an open-source model designed to significantly enhance the factual integrity of AI responses by leveraging its extensive infrastructure. DataGemma combats misinformation by integrating Google’s Data Commons, a massive, interconnected knowledge graph. It utilizes two primary strategies: Retrieval-Augmented Generation, which gathers verified data from the graph before forming answers, and Retrieval-Interleaved Generation, which cross-references generated text against established facts. These methods allow for dynamic dataset curation, moving beyond static training data to ensure outputs are grounded in verifiable reality rather than probabilistic guesses. This initiative is highly relevant to the open data community as it demonstrates how high-quality, structured public datasets can directly improve AI reliability. By releasing these models, Google highlights the critical role of accessible, well-maintained data ecosystems in mitigating AI biases and errors. It underscores that the future of trustworthy artificial intelligence depends heavily on the availability and integration of robust open knowledge graphs, offering a pathway for broader industry adoption of data-driven truthfulness.
Source: androidpolice.comPublished on 2024-09-14
Related news
- Open source AI set to drive new innovation - IT-Online
- Estas son las nuevas herramientas 5.0 para aumentar la productividad de las empresas
- EMBL-EBI Unveils Open Data for Biodiversity, Climate
- New Bureau of Justice Statistics Crime Data Just Released: Violent crime (Rape, Robbery, and Aggravated Assault) Soaring Under Biden-Harris.
- FOIA Friday: Former Virginia Beach staffer’s spending investigated