From the 1 of 11 linked papers with an AI index.
9 papers · 1 filter
EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation
Xinyuan Guan, Feifan Chen, Xinyu Zhan +3
Part-level affordance grounding has advanced the localization of functional object regions associated with elemental actions. Extending this capability to complex tasks calls for c…
Multi-view Hand Reconstruction with a Point-Embedded Transformer
Lixin Yang, Licheng Zhong, Pengxiang Zhu +4
The paper presents POEM, a multi-view hand mesh reconstruction system that embeds static basis points in the multi-view stereo space and uses a transformer to fuse features across…
VFM-Recon: Unlocking Cross-Domain Scene-Level Neural Reconstruction with Scale-Aligned Foundation Priors
Yuhang Ming, Tingkang Xi, Xingrui Yang +4
Scene-level neural volumetric reconstruction from monocular videos remains challenging, especially under severe domain shifts. Although recent advances in vision foundation models…
OmniCam: Unified Multimodal Video Generation via Camera Control
Xiaoda Yang, Jiayang Xu, Kaixuan Luan +9
Camera control, which achieves diverse visual effects by changing camera position and pose, has attracted widespread attention. However, existing methods face challenges such as co…
ChatGarment: Garment Estimation, Generation and Editing via Large Language Models
Siyuan Bian, Chenghao Xu, Yuliang Xiu +5
We introduce ChatGarment, a novel approach that leverages large vision-language models (VLMs) to automate the estimation, generation, and editing of 3D garments from images or text…
COIN: Control-Inpainting Diffusion Prior for Human and Camera Motion Estimation
Jiefeng Li, Ye Yuan, Davis Rempe +5
Estimating global human motion from moving cameras is challenging due to the entanglement of human and camera motions. To mitigate the ambiguity, existing methods leverage learned…