20 citations · 27 across the 6 of their papers we have counts for
8 papers · 1 filter
AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
Difei Gao, Lei Ji, Luowei Zhou +4
Recent research on Large Language Models (LLMs) has led to remarkable advancements in general NLP AI assistants. Some studies have further explored the use of LLMs for planning and…
GroundNLQ @ Ego4D Natural Language Queries Challenge 2023
Zhijian Hou, Lei Ji, Difei Gao +7
In this report, we present our champion solution for Ego4D Natural Language Queries (NLQ) Challenge in CVPR 2023. Essentially, to accurately ground in a video, an effective egocent…
An Efficient COarse-to-fiNE Alignment Framework @ Ego4D Natural Language Queries Challenge 2022
Zhijian Hou, Wanjun Zhong, Lei Ji +6
This technical report describes the CONE approach for Ego4D Natural Language Queries (NLQ) Challenge in ECCV 2022. We leverage our model CONE, an efficient window-centric COarse-to…
Hybrid Reasoning Network for Video-based Commonsense Captioning
Weijiang Yu, Jian Liang, Lei Ji +4
The task of video-based commonsense captioning aims to generate event-wise captions and meanwhile provide multiple commonsense descriptions (e.g., attribute, effect and intention)…
CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
Huaishao Luo, Lei Ji, Ming Zhong +4
Video-text retrieval plays an essential role in multi-modal research and has been widely used in many real-world web applications. The CLIP (Contrastive Language-Image Pre-training…
GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
Chenfei Wu, Lun Huang, Qianxi Zhang +5
Generating videos from text is a challenging task due to its high computational requirements for training and infinite possible answers for evaluation. Existing works typically exp…