How a Cooperative-Override Circuit Suppresses Nash Play in Large Language Models
Researchers identified a cooperative-override circuit in large language models, suppressing Nash play. This finding suggests a word-triggered circuit is responsible for cooperation, rather than a lack of competence. The circuit can be measured, bounded, and controlled. This has implications for understanding and improving the behavior of large language models.
Save an API key to vote.