From the 1 of 6 linked papers with an AI index.
6 papers
Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
Weiquan Lin, Yu Deng, Shiyang Liu +6
This survey systematically reviews how large foundation models provide geometric, semantic, and visual priors for hand‑object interaction tasks such as reconstruction and generatio…
ScaleHP: Scale-Mediated Optimization of Coupled Errors for Metric-Space Hand Pose Estimation
Ruitao Jing, Xingyu Chen, Hongyang Li +3
In this paper, we present ScaleHP, a unified framework that explicitly represents per-instance metric scale to resolve the coupled errors in calibrated camera-space hand pose estim…
HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning
Taowen Wang, Zikang Xie, Bin Yang +13
Humanoid robots promise whole-body interaction in human-centered environments, but scalable policy learning remains difficult because task-level decision-making and whole-body dyna…
SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
Yukai Shi, Weiyu Li, Zihao Wang +4
We propose a decoupled 3D scene generation framework called SceneMaker in this work. Due to the lack of sufficient open-set de-occlusion and pose estimation priors, existing method…
SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
Jinyuan Qu, Hongyang Li, Xingyu Chen +5
In this paper, we present SegDINO3D, a novel Transformer encoder-decoder framework for 3D instance segmentation. As 3D training data is generally not as sufficient as 2D training i…
Detect Anything via Next Point Prediction
Qing Jiang, Junan Huo, Xingyu Chen +6
Object detection has long been dominated by traditional coordinate regression-based models, such as YOLO, DETR, and Grounding DINO. Although recent efforts have attempted to levera…