8 papers · 1 filter
Stereo World Model: Camera-Guided Stereo Video Generation
Yang-Tian Sun, Zehuan Huang, Yifan Niu +4
We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or…
End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
Haoyu Zhang, Wei Zhai, Yuhang Yang +2
Monocular 4D human-object interaction (HOI) reconstruction - recovering a moving human and a manipulated object from a single RGB video - remains challenging due to depth ambiguity…
Event-based Visual Deformation Measurement
Yuliang Wu, Wei Zhai, Yuxin Cui +3
Visual Deformation Measurement (VDM) aims to recover dense deformation fields by tracking surface motion from camera observations. Traditional image-based methods rely on minimal i…
OmniVaT: Single Domain Generalization for Multimodal Visual-Tactile Learning
Liuxiang Qiu, Hui Da, Yuzhen Niu +3
Visual-tactile learning (VTL) enables embodied agents to perceive the physical world by integrating visual (VIS) and tactile (TAC) sensors. However, VTL still suffers from modality…
TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
Guangyi Han, Wei Zhai, Yuhang Yang +2
Hand-object interaction (HOI) is fundamental for humans to express intent. Existing HOI generation research is predominantly confined to fixed grasping patterns, where control is t…
GRACE: Estimating Geometry-level 3D Human-Scene Contact from 2D Images
Chengfeng Wang, Wei Zhai, Yuhang Yang +2
Estimating the geometry level of human-scene contact aims to ground specific contact surface points at 3D human geometries, which provides a spatial prior and bridges the interacti…