7 citations · 11 across the 5 of their papers we have counts for
5 papers
Tencent AVS: A Holistic Ads Video Dataset for Multi-modal Scene Segmentation
Jie Jiang, Zhimin Li, Jiangfeng Xiong +3
Temporal video segmentation and classification have been advanced greatly by public benchmarks in recent years. However, such research still mainly focuses on human actions, failin…
Category-Aware Transformer Network for Better Human-Object Interaction Detection
Leizhen Dong, Zhimin Li, Kunlun Xu +4
Human-Object Interactions (HOI) detection, which aims to localize a human and a relevant object while recognizing their interaction, is crucial for understanding a still image. Rec…
Effective Actor-centric Human-object Interaction Detection
Kunlun Xu, Zhimin Li, Zhijun Zhang +5
While Human-Object Interaction(HOI) Detection has achieved tremendous advances in recent, it still remains challenging due to complex interactions with multiple humans and objects…
Improving Human-Object Interaction Detection via Phrase Learning and Label Composition
Zhimin Li, Cheng Zou, Yu Zhao +2
Human-Object Interaction (HOI) detection is a fundamental task in high-level human-centric scene understanding. We propose PhraseHOI, containing a HOI branch and a novel phrase bra…
Overview of Tencent Multi-modal Ads Video Understanding Challenge
Zhenzhi Wang, Liyu Wu, Zhimin Li +2
Multi-modal Ads Video Understanding Challenge is the first grand challenge aiming to comprehensively understand ads videos. Our challenge includes two tasks: video structuring in t…