activity
20242026
collaborators

7 papers

cs.CV2026

Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation

Siddhant Bansal, Zhifan Zhu, Shashank Tripathi +3

Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing learning-based methods often st…

cs.CV2025

HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding

Jiahe Zhao, Ruibing Hou, Zejie Tian +2

We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). Given a human motion within a 3D scene, HIS-QA requ…

cs.CV2025

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

Jiahe Zhao, Rongkun Zheng, Yi Wang +2

In video Multimodal Large Language Models (video MLLMs), the visual encapsulation process plays a pivotal role in converting video contents into representative tokens for LLM input…

cs.CV2025

unCLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP

Yinqi Li, Jiahe Zhao, Hong Chang +3

Contrastive Language-Image Pre-training (CLIP) has become a foundation model and has been applied to various vision and multimodal tasks. However, recent works indicate that CLIP f…

cs.CV2024

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios

Jie Huang, Ruibing Hou, Jiahe Zhao +2

Human-centric perceptions play a crucial role in real-world applications. While recent human-centric works have achieved impressive progress, these efforts are often constrained to…

cs.CV2024

Clothes-Changing Person Re-Identification with Feasibility-Aware Intermediary Matching

Jiahe Zhao, Ruibing Hou, Hong Chang +4

Current clothes-changing person re-identification (re-id) approaches usually perform retrieval based on clothes-irrelevant features, while neglecting the potential of clothes-relev…