Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models
Researchers propose a new approach to improve the safety of large language models by using 'cunning questions' to train models to detect unusual premises, misleading reasoning, and latent risks. Experiments show that this approach improves robustness to out-of-distribution attacks and strengthens safety fine-tuning.
Save an API key to vote.