Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models

Researchers demonstrated that finetuning large language models can bypass safety alignment strategies, allowing them to reproduce up to 85-90% of copyrighted books, including single verbatim spans exceeding 460 words. This vulnerability affects multiple models from different providers and highlights an industry-wide security issue.

RSS Score 0 9/16/2026, 4:00:00 AM Original Source
Save an API key to vote.