#vision-language models
105 papers · 1 filter
Hearsay: Vision-Language Medical Diagnoses Without an Image
Siddharth Vohra
The paper shows that large vision‑language models generate specific, demographically biased medical diagnoses even when no image is provided, and that this bias appears in structur…
Level, Sharpness, and Corpus: Why Zero-Shot OOD Detector Rankings Do Not Transfer
Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Stephen Gould +1
The paper shows that zero-shot out-of-distribution detector rankings do not reliably transfer across vision-language model deployments and proposes a detector-agnostic wrapper, the…
Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks
Haoyu Zhang, Zhuoxi Wang, Shibo Zheng +7
The paper proposes a guard‑agnostic recovery‑and‑decode module that transcribes encoded or visual text into plain language before applying existing safety classifiers for vision‑la…
MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models
Yitao Zhu, Mengjun Liu, Yingji Fu +2
MedARC is a training-free method that compresses redundant visual tokens in 3D medical images for vision‑language models by scoring token importance with multiple cues and merging…
TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models
Jinkun Zhao, Kui Zhang, Wenjun Wu
The paper introduces Tone‑Pressure Contrastive Decoding (TPCD), which subtracts logits from high‑pressure prompts from those of neutral prompts to reduce commitment bias in vision‑…
TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions
Xinran Liu, Shouqian Shi, Yutong Chen +3
The paper presents TraceCLIP, a training‑free method that extracts patch‑level semantic information from CLIP's CLS attention output to improve zero‑shot dense vision‑language task…