10 papers · 1 filter
GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness
Hui Yang, Wei Sun, Jian Liu +6
Generalized 3D hand-object pose estimation from a single RGB image remains challenging due to the large variations in object appearances and interaction patterns, especially under…
Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control
Nhat Le, Daochang Liu, Anh Nguyen +1
We present MSCoT, a multi-scale, coarse-to-fine model for test-time human motion synthesis and control. Unlike recent approaches that rely on multiple iterative denoising/token-pre…
Diffusion Masked Pretraining for Dynamic Point Cloud
Zhuoyue Zhang, Jihua Zhu, Chaowei Fang +2
Dynamic point cloud pretraining is still dominated by masked reconstruction objectives. However, these objectives inherit two key limitations. Existing methods inject ground-truth…
Mantis: Mamba-native Tuning is Efficient for 3D Point Cloud Foundation Models
Zihao Guo, Jihua Zhu, Jian Liu +1
Pre-trained 3D point cloud foundation models (PFMs) have demonstrated strong transferability across diverse downstream tasks. However, full fine-tuning these models is computationa…
Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility
Yutong Hao, Chen Chen, Ajmal Saeed Mian +2
Diffusion models can generate realistic videos, but existing methods rely on implicitly learning physical reasoning from large-scale text-video datasets, which is costly, difficult…
ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
Zichen Geng, Zeeshan Hayder, Wei Liu +2
3D human reaction generation faces three main challenges:(1) high motion fidelity, (2) real-time inference, and (3) autoregressive adaptability for online scenarios. Existing metho…