7 citations · 15 across the 6 of their papers we have counts for
6 papers
Boosting Weakly-Supervised Image Segmentation via Representation, Transform, and Compensator
Chunyan Wang, Dong Zhang, Rui Yan
Weakly-supervised image segmentation (WSIS) is a critical task in computer vision that relies on image-level class labels. Multi-stage training procedures have been widely used in…
Egocentric Video-Language Pretraining @ Ego4D Challenge 2022
Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan +13
In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Q…
Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022
Kevin Qinghong Lin, Alex Jinpeng Wang, Rui Yan +9
In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for the EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) challenge. Especially,…
Video-Text Pre-training with Learned Regions
Rui Yan, Mike Zheng Shou, Yixiao Ge +4
Video-Text pre-training aims at learning transferable representations from large-scale video-text pairs via aligning the semantics between visual and textual information. State-of-…
Expansion-Squeeze-Excitation Fusion Network for Elderly Activity Recognition
Xiangbo Shu, Jiawen Yang, Rui Yan +1
This work focuses on the task of elderly activity recognition, which is a challenging task due to the existence of individual actions and human-object interactions in elderly activ…
Object-aware Video-language Pre-training for Retrieval
Alex Jinpeng Wang, Yixiao Ge, Guanyu Cai +5
Recently, by introducing large-scale dataset and strong transformer network, video-language pre-training has shown great success especially for retrieval. Yet, existing video-langu…