6 papers
TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models
Zhifang Zhang, Qiqi Tao, Jiaqi Lv +3
The paper introduces TokenSwap, a stealthy backdoor attack on large vision-language models that swaps key textual tokens to corrupt the model's understanding of object relationship…
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
Bingli Wang, Huanze Tang, Haijun Lv +5
In recent years, Multimodal Large Language Models (MLLMs) have achieved remarkable progress on a wide range of multimodal benchmarks. Despite these advances, most existing benchmar…
Do All Individual Layers Help? An Empirical Study of Task-Interfering Layers in Vision-Language Models
Zhiming Liu, Yujie Wei, Lei Feng +5
Current VLMs have demonstrated capabilities across a wide range of multimodal tasks. Typically, in a pretrained VLM, all layers are engaged by default to make predictions on downst…
GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
Zhijie Deng, Chris Yuhao Liu, Zirui Pang +5
Large Language Models (LLMs) have demonstrated strong capabilities in memorizing vast amounts of knowledge across diverse domains. However, the ability to selectively forget specif…
LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models
Zongyu Wu, Yuwei Niu, Hongcheng Gao +12
Large Vision-Language Models (LVLMs) have shown impressive performance in various tasks. However, LVLMs suffer from hallucination, which hinders their adoption in the real world. E…
CNMBERT: A Model for Converting Hanyu Pinyin Abbreviations to Chinese Characters
Zishuo Feng, Feng Cao
The task of converting Hanyu Pinyin abbreviations to Chinese characters is a significant branch within the domain of Chinese Spelling Correction (CSC). It plays an important role i…