works on

From the 1 of 9 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models

Chen Zhong, Xiao An, Zijie Wang +3

Visual inputs in vision-language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bottleneck for efficient inferenc…

cs.CV2026

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model

Chenfeng Wang, Wei He, Xuhan Zhu +10

In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewise benefit from longer latent…

cs.CV2026

SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models

Chen Zhong, Xiao An, Jiaxing Sun +3

Low-level visual perception underpins reliable remote sensing (RS) image analysis, yet current image quality assessment (IQA) methods output uninterpretable scalar scores rather th…

cs.CV2025

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Yiman Zhang, Ziheng Luo, Qiangyu Yan +4

In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with exis…

cs.CV2024

Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs

Kai Han, Jianyuan Guo, Yehui Tang +3

Vision-language large models have achieved remarkable success in various multi-modal tasks, yet applying them to video understanding remains challenging due to the inherent complex…