3 citations · 8 across the 14 of their papers we have counts for
22 papers
STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models
Yichen Guo, Tinghao Wang, Qizhe Zhang +15
Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose substantial computational overhead,…
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
Yichen Guo, Kai Tang, Fenglai Lin +6
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image. Recent…
SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering
Kai Tang, Jinhao You, Bohua Zhang +6
Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answering. However, they remain su…
GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding
Junpeng Ma, Sashuai Zhou, Guanghao Li +9
Video Large Language Models (VLMs) have achieved remarkable success in video understanding, but the significant computational cost from processing dense frames severely limits thei…
SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
Gaole Dai, Menghang Dong, Rongyu Zhang +3
The process through which humans perceive and learn visual representations in dynamic environments is highly complex. From a structural perspective, the human eye decouples the fun…
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
Rongyu Zhang, Menghang Dong, Yuan Zhang +6
Multimodal Large Language Models (MLLMs) excel in understanding complex language and visual data, enabling generalist robotic systems to interpret instructions and perform embodied…