activity
20192026
most citedReading-strategy Inspired Visual Representation Learning for Text-to-Video Retrieval

78 citations · 129 across the 14 of their papers we have counts for

collaborators
Showing cs.CVShow all

14 papers · 1 filter

cs.CV2026

Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain

Daizong Liu, Junhao Dong, Zhiyuan Ma +6

Multimodal large language models (MLLMs) have extended the capability of large language models (LLMs) to process more contextual multimodal information, showing remarkable progress…

cs.CV202278 cited

Reading-strategy Inspired Visual Representation Learning for Text-to-Video Retrieval

Jianfeng Dong, Yabing Wang, Xianke Chen +4

This paper aims for the task of text-to-video retrieval, where given a query in the form of a natural-language sentence, it is asked to retrieve videos which are semantically relev…

cs.CV2022

Unsupervised Temporal Video Grounding with Deep Semantic Clustering

Daizong Liu, Xiaoye Qu, Yinzhen Wang +5

Temporal video grounding (TVG) aims to localize a target segment in a video according to a given sentence query. Though respectable works have made decent achievements in this task…

cs.CV2022

Exploring Motion and Appearance Information for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Pan Zhou +1

This paper addresses temporal sentence grounding. Previous works typically solve this task by learning frame-level video features and align them with the textual information. A maj…

cs.CV2022

Memory-Guided Semantic Learning Network for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Xing Di +3

Temporal sentence grounding (TSG) is crucial and fundamental for video understanding. Although the existing methods train well-designed deep networks with a large amount of data, w…

cs.CV2021

Progressively Guide to Attend: An Iterative Alignment Framework for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Pan Zhou

A key solution to temporal sentence grounding (TSG) exists in how to learn effective alignment between vision and language features extracted from an untrimmed video and a sentence…