Can Data Attribution Filter Out Subliminal Learning? Not Reliably
Researchers evaluated three gradient-based attribution methods for identifying training data responsible for subliminal learning in language models. The methods showed inconsistent results, with EK-FAC providing the most benefit in some settings. The study suggests that gradient-based attribution can be effective but may not be reliable in all cases.
Save an API key to vote.