5 papers
Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain
Daizong Liu, Junhao Dong, Zhiyuan Ma +6
Multimodal large language models (MLLMs) have extended the capability of large language models (LLMs) to process more contextual multimodal information, showing remarkable progress…
Occluding the Solution Space: Planner-Agnostic Adversarial Attacks on Tolerance-Aware Manipulation
Keke Tang, Tianyu Hao, Weilong Peng +5
Adversarial attacks on motion planning are crucial for evaluating and quantifying the intrinsic robustness of robotic manipulation. However, existing approaches are typically limit…
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding
Xiang Fang, Daizong Liu, Wanlong Fang +4
This paper addresses the task of temporal sentence grounding (TSG). Although many respectable works have made decent achievements in this important topic, they severely rely on mas…
Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs
Xiang Fang, Wanlong Fang, Changshuo Wang +4
Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these VLMs are task-specific and…
Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective
Xiang Fang, Zeyu Xiong, Wanlong Fang +7
This paper addresses the challenging task of weakly-supervised video temporal grounding. Existing approaches are generally based on the moment proposal selection framework that uti…