RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue
A new benchmark, RoleBreak, assesses the long-horizon robustness of spoken dialogue models in role-playing scenarios. It evaluates 9 configurations across different paradigms, revealing gaps in both semantic robustness and vocal expressiveness.
Save an API key to vote.