Manipulating Chess-GPT’s World Model

The article provides robust causal evidence that language models like Chess-GPT develop internal "world models" rather than merely memorizing superficial statistical patterns. By using linear probes to identify internal representations of chess board states and player skill levels, the author demonstrates that these latent variables are directly linked to the model’s behavior. This finding challenges the skepticism that large language models lack genuine understanding, suggesting instead that next-token prediction inherently drives the emergence of structured world knowledge. Significant implications arise when the model faces out-of-distribution scenarios, such as randomly initialized chess games where performance typically collapses. The study reveals that this drop occurs because the base model defaults to predicting low-skill moves consistent with its training distribution. However, by intervening in the model’s residual stream to artificially boost its estimated skill level, performance is substantially restored. This ability to modulate behavior through activation editing proves that the model possesses latent strategic capabilities that were previously suppressed by its default predictive tendencies. For the open data and AI research communities, this work is crucial as it offers a practical framework for interpretability and model auditing. It shows that specific, measurable representations within neural networks can be isolated and manipulated to verify internal states, moving beyond black-box analysis. This approach validates the use of small-scale, specialized models to test broader theories about machine cognition, providing open methodologies for demonstrating that self-supervised learning can yield deep, causal understanding of complex domains.

Source: adamkarvonen.github.io
Published on 2024-03-21