collaborators

5 papers

cs.CV2026

Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding

Donghui Feng, Fengxi Zhang, Changsheng Gao +6

Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature…

cs.AI2026

DiffImaginE: Imagine to Verify Entity Types with Diffusion

Feng Zhang, Feiyu Han, Rongxin Yang +11

Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and…

cs.CV2026

EvoCut: Multi-Layer Evolution-Aware Visual Token Compression for Efficient Large Vision-Language Models

Hongyu Lu, Feng Zhang, Wenwei Jin +5

Large vision-language models (LVLMs) achieve strong performance on image and video understanding tasks, but their inference efficiency is constrained by the large number of visual…

cs.CV2026

LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs

Hongyu Lu, Feng Zhang, Wenwei Jin +5

Large vision-language models (LVLMs) achieve strong multimodal understanding, but their inference cost grows rapidly with the number of visual tokens, especially for high-resolutio…

cs.LG2025

Towards Superior Quantization Accuracy: A Layer-sensitive Approach

Feng Zhang, Yanbin Liu, Weihua Li +3

Large Vision and Language Models have exhibited remarkable human-like intelligence in tasks such as natural language comprehension, problem-solving, logical reasoning, and knowledg…