Xeno-Interpretability: Investigating the Alien Minds of LLMs

This research paper introduces the concept of xeno-interpretability, which examines the internal representations of large language models (LLMs) that cannot be adequately expressed in human terms. The study highlights the limitations of current interpretability methods and proposes a new approach to understanding model-native representations.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.