Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
This study found that safety-trained GPT models may not be reducing harm, but rather 'transforming' it by moving discriminatory content from one form to another. This could have implications for the effectiveness of safety protocols in AI agents and the need for a more nuanced approach to evaluating model safety.
Save an API key to vote.