#vision-language models

topicvision-language models

105 papers · 1 filter

cs.CV2026

Hearsay: Vision-Language Medical Diagnoses Without an Image

Siddharth Vohra

The paper shows that large vision‑language models generate specific, demographically biased medical diagnoses even when no image is provided, and that this bias appears in structur…

cs.CV2026

Level, Sharpness, and Corpus: Why Zero-Shot OOD Detector Rankings Do Not Transfer

Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Stephen Gould +1

The paper shows that zero-shot out-of-distribution detector rankings do not reliably transfer across vision-language model deployments and proposes a detector-agnostic wrapper, the…

cs.CR2026

Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks

Haoyu Zhang, Zhuoxi Wang, Shibo Zheng +7

The paper proposes a guard‑agnostic recovery‑and‑decode module that transcribes encoded or visual text into plain language before applying existing safety classifiers for vision‑la…

cs.CV2026

MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models

Yitao Zhu, Mengjun Liu, Yingji Fu +2

MedARC is a training-free method that compresses redundant visual tokens in 3D medical images for vision‑language models by scoring token importance with multiple cues and merging…

cs.CV2026

TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models

Jinkun Zhao, Kui Zhang, Wenjun Wu

The paper introduces Tone‑Pressure Contrastive Decoding (TPCD), which subtracts logits from high‑pressure prompts from those of neutral prompts to reduce commitment bias in vision‑…

cs.CV2026

TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions

Xinran Liu, Shouqian Shi, Yutong Chen +3

The paper presents TraceCLIP, a training‑free method that extracts patch‑level semantic information from CLIP's CLS attention output to improve zero‑shot dense vision‑language task…