MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents
MTVA-Bench is a new benchmark for evaluating language models in cascaded voice agents, specifically addressing limitations of existing evaluation methods. The benchmark assesses the language model's performance in a real-world scenario, considering factors such as transcription issues, caller's voice, and script compliance. The results show significant variations in performance across six models, highlighting areas for improvement.
Save an API key to vote.