most citedMulti-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.AI2025

LMM-Incentive: Large Multimodal Model-based Incentive Design for User-Generated Content in Web 3.0

Jinbo Wen, Jiawen Kang, Linfeng Zhang +5

Web 3.0 represents the next generation of the Internet, which is widely recognized as a decentralized ecosystem that focuses on value expression and data ownership. By leveraging b…

cs.CV2025

Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models

Xuyang Liu, Xiyan Gui, Yuchao Zhang +1

Recent large vision-language models (LVLMs) demonstrate remarkable capabilities in processing extended multi-modal sequences, yet the resulting key-value (KV) cache expansion creat…

cs.CV2025

Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models

Xuyang Liu, Yiyu Wang, Junpeng Ma +1

Video large language models (VideoLLM) excel at video understanding, but face efficiency challenges due to the quadratic complexity of abundant visual tokens. Our systematic analys…

cs.CL2025

Seeing Sarcasm Through Different Eyes: Analyzing Multimodal Sarcasm Perception in Large Vision-Language Models

Junjie Chen, Xuyang Liu, Subin Huang +2

With the advent of large vision-language models (LVLMs) demonstrating increasingly human-like abilities, a pivotal question emerges: do different LVLMs interpret multimodal sarcasm…

cs.CV20241 cited

Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Ting Liu, Liangtao Shi, Richang Hong +3

The vision tokens in multimodal large language models usually exhibit significant spatial and temporal redundancy and take up most of the input tokens, which harms their inference…