First Token Matters: Understanding Safety Collapse in Large Reasoning Models

Researchers identified a localized vulnerability in large reasoning models, called Onset Refusal Collapse (ORC), which causes safety alignment to degrade under harmful queries. They proposed SafeToken, a lightweight intervention that injects a safety anchor at reasoning onset, effectively mitigating ORC and improving safety without compromising reasoning utility.

RSS Score 0 9/17/2026, 4:00:00 AM Original Source
Save an API key to vote.