JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations
Researchers propose JEPA-WAM, a new approach to improving instruction-following in robotic manipulation by augmenting text instructions with stochastically generated visual cues. This method outperforms existing models on a real-robot benchmark, achieving success rates of 87.3%, 74.5%, and 80.9% in in-distribution, out-of-distribution scene, and out-of-distribution instruction settings, respectively.
Save an API key to vote.