activity
20212023
most citedVideo-Text Pre-training with Learned Regions

7 citations · 15 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2023

Boosting Weakly-Supervised Image Segmentation via Representation, Transform, and Compensator

Chunyan Wang, Dong Zhang, Rui Yan

Weakly-supervised image segmentation (WSIS) is a critical task in computer vision that relies on image-level class labels. Multi-stage training procedures have been widely used in…

cs.CV2022

Egocentric Video-Language Pretraining @ Ego4D Challenge 2022

Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan +13

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Q…

cs.CV20221 cited

Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022

Kevin Qinghong Lin, Alex Jinpeng Wang, Rui Yan +9

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for the EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) challenge. Especially,…

cs.CV20217 cited

Video-Text Pre-training with Learned Regions

Rui Yan, Mike Zheng Shou, Yixiao Ge +4

Video-Text pre-training aims at learning transferable representations from large-scale video-text pairs via aligning the semantics between visual and textual information. State-of-…

cs.CV20217 cited

Expansion-Squeeze-Excitation Fusion Network for Elderly Activity Recognition

Xiangbo Shu, Jiawen Yang, Rui Yan +1

This work focuses on the task of elderly activity recognition, which is a challenging task due to the existence of individual actions and human-object interactions in elderly activ…

cs.CV2021

Object-aware Video-language Pre-training for Retrieval

Alex Jinpeng Wang, Yixiao Ge, Guanyu Cai +5

Recently, by introducing large-scale dataset and strong transformer network, video-language pre-training has shown great success especially for retrieval. Yet, existing video-langu…