NewEvery arXiv paper, its researchers & institutions — mapped.
the archive

#vision-language models

105 results
cs.CV2026

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Yao Xiao, Reuben Tan, Zhen Zhu +3

ReToken introduces a single learnable embedding that acts as a retrieval token to select a sparse set of relevant visual tokens from a cached representation, improving vision-langu…

#vision-language models#visual retrieval#sparse token selection#long video processing
cs.AI2026

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Qiushi Sun, Kanzhi Cheng, Yian Wang +20

The paper introduces OSReward, a benchmark for evaluating vision-language model judges that assess computer-using agent trajectories, and presents open reward models (OS‑Shepherd)…

#computer-use agents#vision-language models#reward modeling#benchmark
cs.CV2026

Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs

Jiasheng Li, Zhong Ji, Yan Zhang +1

The paper proposes CaRe, a training‑free method that calibrates compact visual representations before reasoning to keep semantic consistency when reducing visual tokens in large vi…

#visual token reduction#vision-language models#semantic drift#model calibration
cs.CL2026

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

Liangjie Zhao, Jiaqing Lyu, Kexin Tang +5

The paper introduces IllusionReasoning, a benchmark that uses visual illusion images to jointly assess perception and reasoning abilities of large vision‑language models, revealing…

#visual illusions#vision-language models#reasoning evaluation#benchmark
cs.CV2026

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

Haoqing Wang, Xingrun Xing, Wei Xia +2

FaithEyes proposes a multi‑agent framework where a vision‑language model judges its own tool calls to ensure they are useful, improving both accuracy and tool faithfulness on visua…

#vision-language models#tool use#self-judging#reinforcement learning
cs.AI2026

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

Jia Yu, Yan Zhu, Yili He +12

The paper presents EndoCLIP, a vision‑language foundation model for colonoscopy that learns from lesion‑level image‑text pairs extracted from routine colonoscopy reports, achieving…

#colonoscopy#vision-language models#report grounding#image-text retrieval
cs.CV2026

Scaling Vision-Language Models Is Not Enough to Mitigate Bias

Ioannis Sarridis, Ioannis Kompatsiaris, Symeon Papadopoulos

The paper empirically evaluates 194 vision‑language models to see how model size, training data, and architecture affect bias, finding that larger models do not consistently reduce…

#vision-language models#bias mitigation#model scaling#training data quality
cs.RO2026

A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response

Swapnil Saha, Bhuvan Rajanasiriyur Jagadeesha, Karishma Patnaik +1

The paper presents a model‑based systems engineering framework that embeds vision‑language models as coordination agents within a human‑UAV loop for disaster response, demonstratin…

#vision-language models#uav coordination#disaster response#human-autonomy teaming
cs.CV2026

Can Vision-Language Models Reason about AI Edits in Images?

Darsha Udayanga, Pin-Yu Chen, Payel Das +1

The paper explores training vision-language models with reinforcement learning to detect and localize AI-generated image edits, using reasoning traces and a lightweight segmentatio…

#image forgery detection#vision-language models#reinforcement learning#reasoning traces
cs.AI2026

World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models

Xiangcheng Zhang, Yilun Du

The paper introduces World Action Planner, a robot planning system that combines vision‑language models with a multi‑task, pose‑image conditioned world model to generate and iterat…

#action-conditioned world models#vision-language models#robot planning#zero-shot generalization
cs.CV2026

Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction

Lei Yang, Xinze Liu, Dayan Wu +7

The paper introduces a method to detect and correct hallucinated objects in large vision‑language models by identifying a hidden grounding pattern and using a lightweight verifier…

#vision-language models#object hallucination#grounding diagnostics#verifier-guided decoding
cs.AI2026

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

Yang Zhou, Zixuan Huang, Sunzhu Li +10

The paper presents SpatialCLI, a framework that teaches vision-language models to use specialist visual tools for spatial reasoning and then internalize those capabilities, dramati…

#vision-language models#spatial reasoning#tool use#embodied AI
cs.CV2026

MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition

Alex Andonian, Samuel G Rodriques, Andrew D White +1

The paper presents two vision‑language models, OCSRGlyph and MarkushGlyph, that translate images of chemical structures—including single molecules and Markush representations—into…

#optical chemical structure recognition#markush structure parsing#image-to-text translation#stereochemistry handling
cs.AI2026

A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models

Xiangyu Yin, Tora Bodin, Rohan Menon +1

The paper evaluates five direction‑based inference‑time defenses for vision‑language models across multiple architectures, finding that no single method works best for all models a…

#vision-language models#inference-time defenses#direction-based interventions#model robustness
cs.CV2026

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Shawn Li, Wei Yang, Jike Zhong +11

The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…

#visual reasoning#geometric reasoning#jigsaw puzzles#vision-language models
cs.AI2026

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes

Weihang Wang, Kainan Tu, Jielei Zhang +9

The paper presents MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes that evaluates how large vision‑language models handle cultural and background knowledge, an…

#vision-language models#memes#cultural knowledge#benchmark
cs.CV2026

Unifying Adversarially Robust Model Experts in Vision-Language Models

Nguyen Duc Thai, Junhao Dong, Sua Qi Rong +2

The paper introduces CARE, a framework that jointly fine‑tunes multiple adversarially robust vision‑language model experts and merges their knowledge into a single model with compl…

#adversarial robustness#vision-language models#collaborative fine-tuning#embedding alignment
cs.LG2026

Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models

Supratik Bhowal, Subhrajyoti Basu, Aritra Gir Mahanta +1

The paper introduces CoT-Mediate, a framework that edits a medical vision-language model's own generated reasoning to test whether the model's predictions follow that reasoning, re…

#medical imaging#vision-language models#chain-of-thought reasoning#model interpretability
cs.CV2026

MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models

Yitao Zhu, Mengjun Liu, Yingji Fu +2

MedARC is a training-free method that compresses redundant visual tokens in 3D medical images for vision‑language models by scoring token importance with multiple cues and merging…

#3d medical imaging#vision-language models#token compression#adaptive redundancy
cs.CR2026

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

Haoyu Zhang, Zhuoxi Wang, Shibo Zheng +4

The paper proposes a guard‑agnostic recovery‑and‑decode module that transcribes encoded or visual text into plain language before applying existing safety classifiers for vision‑la…

#vision-language models#safety guards#jailbreak attacks#recovery decoding
cs.CV2026

Hearsay: Vision-Language Medical Diagnoses Without an Image

Siddharth Vohra

The paper shows that large vision‑language models generate specific, demographically biased medical diagnoses even when no image is provided, and that this bias appears in structur…

#vision-language models#medical diagnosis#demographic bias#structured output
cs.CV2026

Prior Directions: Why GUI Grounding Gets Locked in the Past

Weile Gong, Zijian Lu, Mingcai Chen +3

The paper investigates how vision-language models can become locked onto outdated textual priors, causing incorrect visual grounding, and identifies recurring latent directions—cal…

#visual grounding#vision-language models#representation analysis#model robustness
cs.CV2026

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Hengyi Xie, Chenfei Yao, Xianjin Wu +7

TurboVLA is a vision-language-action model that directly maps visual observations and language instructions to robot actions, achieving real-time performance (32 Hz) on an RTX 4090…

#vision-language models#robotic manipulation#real-time inference#lightweight architecture
cs.CV2026

Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs

Itbaan Safwan, Ramail Khan, Muhammad Annas Shaikh +1

The paper introduces a multi‑task fine‑tuning approach for small vision‑language models to improve visual question answering on GI endoscopy images, adding grounding and descriptio…

#visual question answering#gastrointestinal endoscopy#multitask learning#vision-language models
← prev1 / 5next →