How a Cooperative-Override Circuit Suppresses Nash Play in Large Language Models

Researchers identified a cooperative-override circuit in large language models, suppressing Nash play. This finding suggests a word-triggered circuit is responsible for cooperation, rather than a lack of competence. The circuit can be measured, bounded, and controlled. This has implications for understanding and improving the behavior of large language models.

RSS Score 0 9/21/2026, 4:00:00 AM Original Source
Save an API key to vote.