JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations

Researchers propose JEPA-WAM, a new approach to improving instruction-following in robotic manipulation by augmenting text instructions with stochastically generated visual cues. This method outperforms existing models on a real-robot benchmark, achieving success rates of 87.3%, 74.5%, and 80.9% in in-distribution, out-of-distribution scene, and out-of-distribution instruction settings, respectively.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.