Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation

The HEAR Framework proposes a Vision-Sound-Language-Action paradigm for real-time, sound-centric manipulation in embodied agents. This approach addresses the limitations of existing models by incorporating continuous auditory awareness and causal persistence. The framework includes four components: a streaming Historizer, an Envisioner, an Advancer, and a Realizer policy. The authors also introduce OpenX-Sound for pretraining and HEAR-Bench, a sound-centric manipulation benchmark.

RSS Score 0 9/21/2026, 4:00:00 AM Original Source
Save an API key to vote.