When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence
Researchers created a simulated benchmark to evaluate the performance of six open vision-language models in diagnosing robot failures. The study found that the models' behavior is often driven by the prompt, not the evidence, and that they struggle to accurately diagnose failures without additional sensor data. The study concludes that the decision to ask for human intervention should be based on measured accuracy and costs, not model confidence.
Save an API key to vote.