collaborators

5 papers

cs.CV2025

Towards Lossless Ultimate Vision Token Compression for VLMs

Dehua Zheng, Mouxiao Huang, Borui Jiang +2

Visual language models encounter challenges in computational efficiency and latency, primarily due to the substantial redundancy in the token representations of high-resolution ima…

cs.CV2025

PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models

Mouxiao Huang, Borui Jiang, Dehua Zheng +3

Multimodal large language models (MLLMs) have achieved strong performance on vision-language tasks, yet often suffer from inefficiencies due to redundant visual tokens. Existing to…

cs.CV2025

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Yiman Zhang, Ziheng Luo, Qiangyu Yan +4

In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with exis…

cs.CV2025

Single Domain Generalization for Few-Shot Counting via Universal Representation Matching

Xianing Chen, Si Huo, Borui Jiang +2

Few-shot counting estimates the number of target objects in an image using only a few annotated exemplars. However, domain shift severely hinders existing methods to generalize to…

cs.CV2025

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Yunxin Li, Zhenyu Liu, Zitao Li +19

Reasoning lies at the heart of intelligence, shaping the ability to make decisions, draw conclusions, and generalize across domains. In artificial intelligence, as systems increasi…