5 papers
OmniVaT: Single Domain Generalization for Multimodal Visual-Tactile Learning
Liuxiang Qiu, Hui Da, Yuzhen Niu +3
Visual-tactile learning (VTL) enables embodied agents to perceive the physical world by integrating visual (VIS) and tactile (TAC) sensors. However, VTL still suffers from modality…
TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
Guangyi Han, Wei Zhai, Yuhang Yang +2
Hand-object interaction (HOI) is fundamental for humans to express intent. Existing HOI generation research is predominantly confined to fixed grasping patterns, where control is t…
GRACE: Estimating Geometry-level 3D Human-Scene Contact from 2D Images
Chengfeng Wang, Wei Zhai, Yuhang Yang +2
Estimating the geometry level of human-scene contact aims to ground specific contact surface points at 3D human geometries, which provides a spatial prior and bridges the interacti…
VideoGen-Eval: Agent-based System for Video Generation Evaluation
Yuhang Yang, Ke Fan, Shangkun Sun +7
The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot sho…
SIGMAN:Scaling 3D Human Gaussian Generation with Millions of Assets
Yuhang Yang, Fengqi Liu, Yixing Lu +8
3D human digitization has long been a highly pursued yet challenging task. Existing methods aim to generate high-quality 3D digital humans from single or multiple views, but remain…