activity
20212025
most citedLearning to Decompose Visual Features with Latent Textual Prompts

8 citations · 42 across the 24 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2023

ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representation

Yangyi Chen, Xingyao Wang, Manling Li +2

State-of-the-art vision-language models (VLMs) still have limited performance in structural knowledge extraction, such as relations between objects. In this work, we present ViStru…

cs.CV2022★ 2 cited

Video Event Extraction via Tracking Visual States of Arguments

Guang Yang, Manling Li, Jiajie Zhang +3

Video event extraction aims to detect salient events from a video and identify the arguments for each event as well as their semantic roles. Existing methods focus on capturing the…

cs.CV2022★ 8 cited

Learning to Decompose Visual Features with Latent Textual Prompts

Feng Wang, Manling Li, Xudong Lin +3

Recent advances in pre-training vision-language models like CLIP have shown great potential in learning transferable visual representations. Nonetheless, for downstream inference,…

cs.CV2022★ 1 cited

Towards Fast Adaptation of Pretrained Contrastive Models for Multi-channel Video-Language Retrieval

Xudong Lin, Simran Tiwari, Shiyuan Huang +4

Multi-channel video-language retrieval require models to understand information from different channels (e.g. videoquestion, videospeech) to correctly link a video with a tex…

cs.CV2021★ 3 cited

Joint Multimedia Event Extraction from Video and Article

Brian Chen, Xudong Lin, Christopher Thomas +5

Visual and textual modalities contribute complementary information about events described in multimedia documents. Videos contain rich dynamics and detailed unfoldings of events, w…