FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models
This paper introduces FOCAL-VLA, a framework that enhances vision-language-action (VLA) models through subtask-guided geometry distillation and implicit world modeling. This improves VLA models' ability to learn spatial and temporal understanding, leading to better performance in robotic manipulation tasks.
Save an API key to vote.