activity
20182022
most citedXGPT: Cross-modal Generative Pre-Training for Image Captioning

20 citations · 27 across the 6 of their papers we have counts for

collaborators

11 papers

cs.CV20221 cited

An Efficient COarse-to-fiNE Alignment Framework @ Ego4D Natural Language Queries Challenge 2022

Zhijian Hou, Wanjun Zhong, Lei Ji +6

This technical report describes the CONE approach for Ego4D Natural Language Queries (NLQ) Challenge in ECCV 2022. We leverage our model CONE, an efficient window-centric COarse-to…

cs.CV20212 cited

Hybrid Reasoning Network for Video-based Commonsense Captioning

Weijiang Yu, Jian Liang, Lei Ji +4

The task of video-based commonsense captioning aims to generate event-wise captions and meanwhile provide multiple commonsense descriptions (e.g., attribute, effect and intention)…

cs.CL2021

GEM: A General Evaluation Benchmark for Multimodal Tasks

Lin Su, Nan Duan, Edward Cui +7

In this paper, we present GEM as a General Evaluation benchmark for Multimodal tasks. Different from existing datasets such as GLUE, SuperGLUE, XGLUE and XTREME that mainly focus o…

cs.CV2021

CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Huaishao Luo, Lei Ji, Ming Zhong +4

Video-text retrieval plays an essential role in multi-modal research and has been widely used in many real-world web applications. The CLIP (Contrastive Language-Image Pre-training…

cs.CV2021

GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Chenfei Wu, Lun Huang, Qianxi Zhang +5

Generating videos from text is a challenging task due to its high computational requirements for training and infinite possible answers for evaluation. Existing works typically exp…

cs.CL20204 cited

GRACE: Gradient Harmonized and Cascaded Labeling for Aspect-based Sentiment Analysis

Huaishao Luo, Lei Ji, Tianrui Li +2

In this paper, we focus on the imbalance issue, which is rarely studied in aspect term extraction and aspect sentiment classification when regarding them as sequence labeling tasks…