The paper presents ViTacWorld, a framework for scalable contact-rich robot manipulation using a visuo-tactile world model.
First framework for robot visuo-tactile-action trajectory generation and policy evaluation
Before reading this…
Applications
To understand this paper, make sure you know these concepts first:
Contact-rich robot manipulation requires physical interaction cues that are often invisible to cameras, making tactile sensing essential for robust control. However, scaling visuo-tactile robot learning remains difficult because real tactile interaction data are expensive to collect, hardware-dependent, and limited in task and scene diversity. We present ViTacWorld, an action-conditioned visuo-tactile world model for scalable contact-rich robot manipulation. ViTacWorld leverages public real tactile datasets and a constructed simulation environment to scale visuo-tactile-action data, exploiting the fact that tactile signals are directly grounded in physical contact and can exhibit a smaller simulation-to-real gap than purely visual observations. The model is first pretrained with large-scale real and simulated visuo-tactile trajectories, and then finetuned with real-world policy rollouts to better match downstream manipulation behaviors. Given robot actions, ViTacWorld predicts temporally aligned visual observations and tactile feedback, enabling visuo-tactile-action rollout generation. To the best of our knowledge, ViTacWorld is the first framework that uses a world model for robot visuo-tactile-action trajectory generation and policy evaluation. It serves two roles: synthesizing rollouts to improve downstream tactile policies, and evaluating policies by predicting action-conditioned visuo-tactile outcomes under controlled action sequences. Experiments on contact-rich manipulation tasks show that ViTacWorld generates physically meaningful rollouts, improves policy performance through scalable data augmentation, and enables action-conditioned policy evaluation. Project page: https://vitacworld.github.io/
VTLoc: Learning-based Tactile Contact Localization in Visual Point Clouds
The paper proposes VTLoc, a framework that aligns and refines contact points bet…
BayesContact: Uncertain Pose Estimation via Visuo-Tactile Proposals and Simulation-based Inference
The paper introduces BayesContact, a framework for visuo-tactile pose estimation…
Robot-Factored World Models via Robot Rendering
This paper proposes robot-factored world models for action-conditioned video pre…
Data Pyramid for Embodied Manipulation
This paper organizes embodied data sources for multimodal foundation models into…
DriftWorld: Fast World Modeling through Drifting
This paper introduces DriftWorld, an action-conditioned world model based on dri…
Data and Learning Where it Matters for Contact-Rich Manipulation
The paper proposes an automated data-collection scheme for contact-rich tasks us…
Design and Human Evaluation of Tactile Withdrawal Reflexes for a Skin-Covered Robot Arm
This paper presents a complete pipeline for artificial nociception in a robotic…
RoboTTT: Context Scaling for Robot Policies
This paper introduces Test-Time-Training Robot Policies (RoboTTT), a robot model…