Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks
Researchers demonstrated a vulnerability in self-modifying AI coding agents, allowing an adversary to poison benchmarks and induce the agent to write vulnerable code. This attack has implications for AI agent security, particularly when agents generate new versions of themselves.
Save an API key to vote.