3 papers
cs.CV2026
Occlusion-Robust Multi-Object Decoupling for Physics-Based Robotic Interaction
Xin Dong, Lihan Zhang, Tianru Dai +2
We propose a mask-free method for lossless multi-object 3D reconstruction from sparse and occluded real-world views, enabling physically plausible robotic interaction via Material…
cs.RO2026
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
Chubin Zhang, Jianan Wang, Zifeng Gao +5
Generalist Vision-Language-Action models remain constrained by the scarcity of robotic data relative to the abundance of human video demonstrations. Existing Latent Action Models a…
cs.CV2026
CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction
Xin Dong, Weijian Deng, Lihan Zhang +3
Reconstructing interaction-ready 3D worlds is essential for physical simulation, virtual reality, robotics, and autonomous driving. However, existing methods mainly optimize static…