2 citations · 6 across the 16 of their papers we have counts for
10 papers · 1 filter
Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain
Daizong Liu, Junhao Dong, Zhiyuan Ma +6
Multimodal large language models (MLLMs) have extended the capability of large language models (LLMs) to process more contextual multimodal information, showing remarkable progress…
Rethinking Video-Language Model from the Language Input Perspective
Xiang Fang, Wanlong Fang, Changshuo Wang +2
Driven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between videos and texts. Although…
Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs
Xiang Fang, Wanlong Fang, Changshuo Wang +4
Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these VLMs are task-specific and…
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
Siyuan Huang, Xiaoye Qu, Yafu Li +6
While autoregressive Large Vision-Language Models (LVLMs) demonstrate remarkable proficiency in multimodal tasks, they face a "Visual Signal Dilution" phenomenon, where the accumul…
VideoSSR: Video Self-Supervised Reinforcement Learning
Zefeng He, Xiaoye Qu, Yafu Li +3
Reinforcement Learning with Verifiable Rewards (RLVR) has substantially advanced the video understanding capabilities of Multimodal Large Language Models (MLLMs). However, the rapi…
Dual Learning with Dynamic Knowledge Distillation and Soft Alignment for Partially Relevant Video Retrieval
Jianfeng Dong, Lei Huang, Daizong Liu +5
Almost all previous text-to-video retrieval works ideally assume that videos are pre-trimmed with short durations containing solely text-related content. However, in practice, vide…