Xeno-Interpretability: Investigating the Alien Minds of LLMs
This research paper introduces the concept of xeno-interpretability, which examines the internal representations of large language models (LLMs) that cannot be adequately expressed in human terms. The study highlights the limitations of current interpretability methods and proposes a new approach to understanding model-native representations.
Save an API key to vote.