3 papers
cs.CV2026
VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation
Tianxiao Chen, Hanmo Chen, Huajin Chen +3
Modern video generation models can synthesize visually compelling and temporally coherent clips, yet controlling their physical behavior remains difficult with standard text and im…
cs.RO2025
VTAO-BiManip: Masked Visual-Tactile-Action Pre-training with Object Understanding for Bimanual Dexterous Manipulation
Zhengnan Sun, Zhaotai Shi, Jiayin Chen +4
Bimanual dexterous manipulation remains significant challenges in robotics due to the high DoFs of each hand and their coordination. Existing single-hand manipulation techniques of…
cs.CV2025
OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
Zhenhao Zhang, Ye Shi, Lingxiao Yang +3
Understanding and synthesizing realistic 3D hand-object interactions (HOI) is critical for applications ranging from immersive AR/VR to dexterous robotics. Existing methods struggl…