activity
20122025
most citedGoing Deeper with Convolutions

1.4k citations · 1.9k across the 24 of their papers we have counts for

collaborators
Showing 2022 · cs.CVShow all

9 papers · 2 filters

cs.CV2022★ 2 cited

Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning

Yatai Ji, Rongcheng Tu, Jie Jiang +6

Cross-modal alignment is essential for vision-language pre-training (VLP) models to learn the correct corresponding information across different modalities. For this purpose, inspi…

cs.CV2022

Egocentric Video-Language Pretraining @ Ego4D Challenge 2022

Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan +13

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Q…

cs.CV2022★ 1 cited

Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022

Kevin Qinghong Lin, Alex Jinpeng Wang, Rui Yan +9

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for the EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) challenge. Especially,…

cs.CV2022

Boosting Multi-Modal E-commerce Attribute Value Extraction via Unified Learning Scheme and Dynamic Range Minimization

Mengyin Liu, Chao Zhu, Hongyu Gao +4

With the prosperity of e-commerce industry, various modalities, e.g., vision and language, are utilized to describe product items. It is an enormous challenge to understand such di…

cs.CV2022

Towards Generalizable Person Re-identification with a Bi-stream Generative Model

Xin Xu, Wei Liu, Zheng Wang +2

Generalizable person re-identification (re-ID) has attracted growing attention due to its powerful adaptation capability in the unseen data domain. However, existing solutions ofte…

cs.CV2022★ 11 cited

Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning

Li Yang, Yan Xu, Chunfeng Yuan +3

Visual grounding is a task to locate the target indicated by a natural language expression. Existing methods extend the generic object detection framework to this problem. They bas…