GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation
Researchers introduce GUARD, a method to guide large reasoning models to forget sensitive information by converting unsafe disclosures into safe-exit trajectories. They also propose a new metric, Natural Forgetting Reasoning Score (NFRS), to evaluate the quality of forgetting. Experiments show that GUARD reduces unsafe disclosures while preserving reasoning utility in two widely adopted distilled LRMs.
Save an API key to vote.