#vision-language models
105 resultsReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Yao Xiao, Reuben Tan, Zhen Zhu +3
ReToken introduces a single learnable embedding that acts as a retrieval token to select a sparse set of relevant visual tokens from a cached representation, improving vision-langu…
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Qiushi Sun, Kanzhi Cheng, Yian Wang +20
The paper introduces OSReward, a benchmark for evaluating vision-language model judges that assess computer-using agent trajectories, and presents open reward models (OS‑Shepherd)…
Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs
Jiasheng Li, Zhong Ji, Yan Zhang +1
The paper proposes CaRe, a training‑free method that calibrates compact visual representations before reasoning to keep semantic consistency when reducing visual tokens in large vi…
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
Liangjie Zhao, Jiaqing Lyu, Kexin Tang +5
The paper introduces IllusionReasoning, a benchmark that uses visual illusion images to jointly assess perception and reasoning abilities of large vision‑language models, revealing…
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
Haoqing Wang, Xingrun Xing, Wei Xia +2
FaithEyes proposes a multi‑agent framework where a vision‑language model judges its own tool calls to ensure they are useful, improving both accuracy and tool faithfulness on visua…
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
Jia Yu, Yan Zhu, Yili He +12
The paper presents EndoCLIP, a vision‑language foundation model for colonoscopy that learns from lesion‑level image‑text pairs extracted from routine colonoscopy reports, achieving…
Scaling Vision-Language Models Is Not Enough to Mitigate Bias
Ioannis Sarridis, Ioannis Kompatsiaris, Symeon Papadopoulos
The paper empirically evaluates 194 vision‑language models to see how model size, training data, and architecture affect bias, finding that larger models do not consistently reduce…
A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response
Swapnil Saha, Bhuvan Rajanasiriyur Jagadeesha, Karishma Patnaik +1
The paper presents a model‑based systems engineering framework that embeds vision‑language models as coordination agents within a human‑UAV loop for disaster response, demonstratin…
Can Vision-Language Models Reason about AI Edits in Images?
Darsha Udayanga, Pin-Yu Chen, Payel Das +1
The paper explores training vision-language models with reinforcement learning to detect and localize AI-generated image edits, using reasoning traces and a lightweight segmentatio…
World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
Xiangcheng Zhang, Yilun Du
The paper introduces World Action Planner, a robot planning system that combines vision‑language models with a multi‑task, pose‑image conditioned world model to generate and iterat…
Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction
Lei Yang, Xinze Liu, Dayan Wu +7
The paper introduces a method to detect and correct hallucinated objects in large vision‑language models by identifying a hidden grounding pattern and using a lightweight verifier…
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
Yang Zhou, Zixuan Huang, Sunzhu Li +10
The paper presents SpatialCLI, a framework that teaches vision-language models to use specialist visual tools for spatial reasoning and then internalize those capabilities, dramati…
MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition
Alex Andonian, Samuel G Rodriques, Andrew D White +1
The paper presents two vision‑language models, OCSRGlyph and MarkushGlyph, that translate images of chemical structures—including single molecules and Markush representations—into…
A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models
Xiangyu Yin, Tora Bodin, Rohan Menon +1
The paper evaluates five direction‑based inference‑time defenses for vision‑language models across multiple architectures, finding that no single method works best for all models a…
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Shawn Li, Wei Yang, Jike Zhong +11
The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…
MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes
Weihang Wang, Kainan Tu, Jielei Zhang +9
The paper presents MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes that evaluates how large vision‑language models handle cultural and background knowledge, an…
Unifying Adversarially Robust Model Experts in Vision-Language Models
Nguyen Duc Thai, Junhao Dong, Sua Qi Rong +2
The paper introduces CARE, a framework that jointly fine‑tunes multiple adversarially robust vision‑language model experts and merges their knowledge into a single model with compl…
Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models
Supratik Bhowal, Subhrajyoti Basu, Aritra Gir Mahanta +1
The paper introduces CoT-Mediate, a framework that edits a medical vision-language model's own generated reasoning to test whether the model's predictions follow that reasoning, re…
MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models
Yitao Zhu, Mengjun Liu, Yingji Fu +2
MedARC is a training-free method that compresses redundant visual tokens in 3D medical images for vision‑language models by scoring token importance with multiple cues and merging…
Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks
Haoyu Zhang, Zhuoxi Wang, Shibo Zheng +4
The paper proposes a guard‑agnostic recovery‑and‑decode module that transcribes encoded or visual text into plain language before applying existing safety classifiers for vision‑la…
Hearsay: Vision-Language Medical Diagnoses Without an Image
Siddharth Vohra
The paper shows that large vision‑language models generate specific, demographically biased medical diagnoses even when no image is provided, and that this bias appears in structur…
Prior Directions: Why GUI Grounding Gets Locked in the Past
Weile Gong, Zijian Lu, Mingcai Chen +3
The paper investigates how vision-language models can become locked onto outdated textual priors, causing incorrect visual grounding, and identifies recurring latent directions—cal…
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Hengyi Xie, Chenfei Yao, Xianjin Wu +7
TurboVLA is a vision-language-action model that directly maps visual observations and language instructions to robot actions, achieving real-time performance (32 Hz) on an RTX 4090…
Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs
Itbaan Safwan, Ramail Khan, Muhammad Annas Shaikh +1
The paper introduces a multi‑task fine‑tuning approach for small vision‑language models to improve visual question answering on GI endoscopy images, adding grounding and descriptio…