PACT-WAM: Predicting Actions and Visual Foresight with Compact Temporal Encoding for Robot Manipulation
A new model, PACT-WAM, is introduced for robot manipulation tasks, using temporal context and visual foresight to predict actions and their consequences. The model reduces processing costs by using a compact temporal encoding and hierarchical history representation. It achieves high success rates on various tasks and supports a test-time enhancement through a vision-language model component.
Save an API key to vote.