15 citations · 29 across the 2 of their papers we have counts for
5 papers · 1 filter
AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any Scenario
Yuhan Li, Hao Zhou, Wenxiang Shang +3
While image-based virtual try-on has made significant strides, emerging approaches still fall short of delivering high-fidelity and robust fitting images across various scenarios,…
Towards Diverse Temporal Grounding under Single Positive Labels
Hao Zhou, Chongyang Zhang, Yanjun Chen +1
Temporal grounding aims to retrieve moments of the described event within an untrimmed video by a language query. Typically, existing methods assume annotations are precise and uni…
Embracing Uncertainty: Decoupling and De-bias for Robust Temporal Grounding
Hao Zhou, Chongyang Zhang, Yan Luo +2
Temporal grounding aims to localize temporal boundaries within untrimmed videos by language queries, but it faces the challenge of two types of inevitable human uncertainties: quer…
Where, What, Whether: Multi-modal Learning Meets Pedestrian Detection
Yan Luo, Chongyang Zhang, Muming Zhao +2
Pedestrian detection benefits greatly from deep convolutional neural networks (CNNs). However, it is inherently hard for CNNs to handle situations in the presence of occlusion and…
Visual Relationship Detection with Relative Location Mining
Hao Zhou, Chongyang Zhang, Chuanping Hu
Visual relationship detection, as a challenging task used to find and distinguish the interactions between object pairs in one image, has received much attention recently. In this…