Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models
A new hierarchical long-horizon vision-language-action architecture with an explicit language-memory module is proposed to improve the success rate and robustness of VLA models on complex tasks. The architecture decouples the system into a high-level VLM and a low-level VLA, enabling persistent temporal tracking and dynamic correction during long-horizon execution.
Save an API key to vote.