AIs get worse at answering simple questions as they get bigger

Larger language models demonstrate improved performance on complex tasks through scaling and fine-tuning, yet they become less reliable for simple queries. As models gain confidence in handling intricate challenges, their tendency to provide incorrect answers for basic questions increases, revealing a critical degradation in foundational accuracy. This paradox suggests that enhanced capability does not necessarily translate to consistent overall reliability. The findings underscore a significant mismatch between user perception and actual AI capability. By presenting systems as omniscient, developers inadvertently foster an unjustified overreliance among users who trust these tools more than warranted. Since models cannot accurately identify their own knowledge limits, they risk providing false information, posing serious dangers in contexts requiring precise factual correctness. This research is vital for open data initiatives, which increasingly rely on automated text processing. If foundational models struggle with basic arithmetic or facts, datasets cleaned or generated by them may contain hidden errors. Open data projects must therefore implement rigorous human validation to ensure integrity, recognizing that algorithmic advancement does not eliminate the need for transparent, verified data sources.

Source: newscientist.com
Published on 2024-09-26