4 papers
EvoCut: Multi-Layer Evolution-Aware Visual Token Compression for Efficient Large Vision-Language Models
Hongyu Lu, Feng Zhang, Wenwei Jin +5
Large vision-language models (LVLMs) achieve strong performance on image and video understanding tasks, but their inference efficiency is constrained by the large number of visual…
LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs
Hongyu Lu, Feng Zhang, Wenwei Jin +5
Large vision-language models (LVLMs) achieve strong multimodal understanding, but their inference cost grows rapidly with the number of visual tokens, especially for high-resolutio…
Sparsity and Out-of-Distribution Generalization
Scott Aaronson, Lin Lin Lee, Jiawei Li
Explaining out-of-distribution generalization has been a central problem in epistemology since Goodman's "grue" puzzle in 1946. Today it's a central problem in machine learning, in…
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Guoxuan Chen, Han Shi, Jiawei Li +7
Large Language Models (LLMs) have exhibited exceptional performance across a spectrum of natural language processing tasks. However, their substantial sizes pose considerable chall…