5 papers
CSPO: Alleviating Reward Ambiguity for Structured Table-to-LaTeX Generation
Yunfan Yang, Cuiling Lan, Jitao Sang +1
Tables contain rich structured information, yet when stored as images their contents remain "locked" within pixels. Converting table images into LaTeX code enables faithful digitiz…
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models
Jiaming Zhang, Junhong Ye, Xingjun Ma +5
Due to their multimodal capabilities, Vision-Language Models (VLMs) have found numerous impactful applications in real-world scenarios. However, recent studies have revealed that V…
Mind with Eyes: from Language Reasoning to Multimodal Reasoning
Zhiyu Lin, Yifei Gao, Xian Zhao +2
Language models have recently advanced into the realm of reasoning, yet it is through multimodal reasoning that we can fully unlock the potential to achieve more comprehensive, hum…
Debiased Prompt Tuning in Vision-Language Model without Annotations
Chaoquan Jiang, Yunfan Yang, Rui Hu +1
Prompt tuning of Vision-Language Models (VLMs) such as CLIP, has demonstrated the ability to rapidly adapt to various downstream tasks. However, recent studies indicate that tuned…
Debiasing Vison-Language Models with Text-Only Training
Yunfan Yang, Chaoquan Jiang, Zhiyu Lin +3
Pre-trained vision-language models (VLMs), such as CLIP, have exhibited remarkable performance across various downstream tasks by aligning text and images in a unified embedding sp…