Verbalizing Subliminal Learning Effects Using Text Optimization
Researchers developed a method called SALVE to detect subliminal learning effects in AI models by using text optimization to recover legible prompts from a distillation dataset. This method can identify traits from the teacher model not encoded in the dataset, and can be used to detect subliminal learning effects in various settings.
Save an API key to vote.