RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

A new benchmark, RoleBreak, assesses the long-horizon robustness of spoken dialogue models in role-playing scenarios. It evaluates 9 configurations across different paradigms, revealing gaps in both semantic robustness and vocal expressiveness.

RSS Score 0 9/16/2026, 4:00:00 AM Original Source
Save an API key to vote.