7 papers
Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation
Siddhant Bansal, Zhifan Zhu, Shashank Tripathi +3
Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing learning-based methods often st…
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
Jiahe Zhao, Ruibing Hou, Zejie Tian +2
We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). Given a human motion within a 3D scene, HIS-QA requ…
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
Jiahe Zhao, Rongkun Zheng, Yi Wang +2
In video Multimodal Large Language Models (video MLLMs), the visual encapsulation process plays a pivotal role in converting video contents into representative tokens for LLM input…
unCLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
Yinqi Li, Jiahe Zhao, Hong Chang +3
Contrastive Language-Image Pre-training (CLIP) has become a foundation model and has been applied to various vision and multimodal tasks. However, recent works indicate that CLIP f…
RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios
Jie Huang, Ruibing Hou, Jiahe Zhao +2
Human-centric perceptions play a crucial role in real-world applications. While recent human-centric works have achieved impressive progress, these efforts are often constrained to…
Clothes-Changing Person Re-Identification with Feasibility-Aware Intermediary Matching
Jiahe Zhao, Ruibing Hou, Hong Chang +4
Current clothes-changing person re-identification (re-id) approaches usually perform retrieval based on clothes-irrelevant features, while neglecting the potential of clothes-relev…