Can Data Attribution Filter Out Subliminal Learning? Not Reliably

Researchers evaluated three gradient-based attribution methods for identifying training data responsible for subliminal learning in language models. The methods showed inconsistent results, with EK-FAC providing the most benefit in some settings. The study suggests that gradient-based attribution can be effective but may not be reliable in all cases.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.