World Modeling in Transformers
This research explores the concept of world modeling in transformer models, specifically in the context of a transformer trained on random walks through Manhattan. The study reveals that the model represents its environment, tracks its position, and uses a goal compass to navigate, contradicting previous claims of incoherent internal maps. The findings suggest that world-modeling capacities emerge at different stages of training and propose mechanistic indicators for comparing models.
Save an API key to vote.