78 citations · 129 across the 14 of their papers we have counts for
14 papers · 1 filter
Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain
Daizong Liu, Junhao Dong, Zhiyuan Ma +6
Multimodal large language models (MLLMs) have extended the capability of large language models (LLMs) to process more contextual multimodal information, showing remarkable progress…
Reading-strategy Inspired Visual Representation Learning for Text-to-Video Retrieval
Jianfeng Dong, Yabing Wang, Xianke Chen +4
This paper aims for the task of text-to-video retrieval, where given a query in the form of a natural-language sentence, it is asked to retrieve videos which are semantically relev…
Unsupervised Temporal Video Grounding with Deep Semantic Clustering
Daizong Liu, Xiaoye Qu, Yinzhen Wang +5
Temporal video grounding (TVG) aims to localize a target segment in a video according to a given sentence query. Though respectable works have made decent achievements in this task…
Exploring Motion and Appearance Information for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Pan Zhou +1
This paper addresses temporal sentence grounding. Previous works typically solve this task by learning frame-level video features and align them with the textual information. A maj…
Memory-Guided Semantic Learning Network for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Xing Di +3
Temporal sentence grounding (TSG) is crucial and fundamental for video understanding. Although the existing methods train well-designed deep networks with a large amount of data, w…
Progressively Guide to Attend: An Iterative Alignment Framework for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Pan Zhou
A key solution to temporal sentence grounding (TSG) exists in how to learn effective alignment between vision and language features extracted from an untrimmed video and a sentence…