Bypassing the Rationale: Causal Auditing of Implicit Reasoning in Language Models
Researchers introduced a method to audit the causal faithfulness of chain-of-thought (CoT) prompting in language models. They found that CoT's influence is often localized and can be bypassed in certain regimes, and that model tuning and architecture can affect CoT's impact. This affects developers who rely on CoT as a transparency mechanism, as it may not always reflect the model's internal computation.
Save an API key to vote.