15 papers
GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness
Hui Yang, Wei Sun, Jian Liu +6
Generalized 3D hand-object pose estimation from a single RGB image remains challenging due to the large variations in object appearances and interaction patterns, especially under…
Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control
Nhat Le, Daochang Liu, Anh Nguyen +1
We present MSCoT, a multi-scale, coarse-to-fine model for test-time human motion synthesis and control. Unlike recent approaches that rely on multiple iterative denoising/token-pre…
Diffusion Masked Pretraining for Dynamic Point Cloud
Zhuoyue Zhang, Jihua Zhu, Chaowei Fang +2
Dynamic point cloud pretraining is still dominated by masked reconstruction objectives. However, these objectives inherit two key limitations. Existing methods inject ground-truth…
Mantis: Mamba-native Tuning is Efficient for 3D Point Cloud Foundation Models
Zihao Guo, Jihua Zhu, Jian Liu +1
Pre-trained 3D point cloud foundation models (PFMs) have demonstrated strong transferability across diverse downstream tasks. However, full fine-tuning these models is computationa…
Causal Reinforcement Learning for Complex Card Games: A Magic The Gathering Benchmark
Cristiano da Costa Cunha, Ajmal Mian, Tim French +1
Causal reinforcement learning (RL) lacks benchmarks for complex systems that combine sequential decision making, hidden information, large masked action spaces, and explicit causal…
Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility
Yutong Hao, Chen Chen, Ajmal Saeed Mian +2
Diffusion models can generate realistic videos, but existing methods rely on implicitly learning physical reasoning from large-scale text-video datasets, which is costly, difficult…