Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models
Researchers demonstrated that finetuning large language models can bypass safety alignment strategies, allowing them to reproduce up to 85-90% of copyrighted books, including single verbatim spans exceeding 460 words. This vulnerability affects multiple models from different providers and highlights an industry-wide security issue.
Save an API key to vote.