collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning

Haocong He, Chenfei Liao, Zichen Wen +13

Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the emerging visual modality of p…

cs.CV2026

TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition

Junyuan Zhang, Bin Wang, Qintong Zhang +13

Table recognition (TR) aims to transform table images into semi-structured representations such as HTML or Markdown. As a core component of document parsing, TR has long relied on…

cs.CV2026

DRUPI: Dataset Reduction Using Privileged Information

Shaobo Wang, Youxin Jiang, Tianle Niu +9

Dataset Condensation (DC) seeks to select or distill samples from large datasets into smaller subsets while preserving performance on target tasks. Existing methods primarily focus…

cs.CV2026

Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving

Minhao Xiong, Zichen Wen, Zhuangcheng Gu +9

Vision-Language Models (VLMs) have emerged as a promising paradigm in autonomous driving (AD), providing a unified framework for perception and decision-making. However, their real…

cs.CV2025

IPCV: Information-Preserving Compression for MLLM Visual Encoders

Yuan Chen, Zichen Wen, Yuzhou Wu +6

Multimodal Large Language Models (MLLMs) deliver strong vision-language performance but at high computational cost, driven by numerous visual tokens processed by the Vision Transfo…

cs.CV2025

VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction

Shaobo Wang, Tianle Niu, Runkang Yang +6

The scalability of video understanding models is increasingly limited by the prohibitive storage and computational costs of large-scale video datasets. While data synthesis has imp…