collaborators

6 papers

cs.CV2025

FastVID: Dynamic Density Pruning for Fast Video Large Language Models

Leqi Shen, Guoqiang Gong, Tao He +4

Video Large Language Models have demonstrated strong video understanding capabilities, yet their practical deployment is hindered by substantial inference costs caused by redundant…

cs.CL2025

SC: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models

Tao He, Guang Huang, Yu Yang +5

Large language models (LLMs) exhibit remarkable reasoning capabilities across diverse downstream tasks. However, their autoregressive nature leads to substantial inference latency,…

cs.CV2025

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval

Leqi Shen, Guoqiang Gong, Tianxiang Hao +6

The parameter-efficient adaptation of the image-text pretraining model CLIP for video-text retrieval is a prominent area of research. While CLIP is focused on image-level vision-la…

cs.CV2025

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models

Fengyuan Sun, Leqi Shen, Hui Chen +3

Video Large Language Models (Video LLMs) have achieved remarkable results in video understanding tasks. However, they often suffer from heavy computational overhead due to the larg…

cs.IR2025

Modality Reliability Guided Multimodal Recommendation

Xue Dong, Xuemeng Song, Na Zheng +2

Multimodal recommendation faces an issue of the performance degradation that the uni-modal recommendation sometimes achieves the better performance. A possible reason is that the u…

cs.CV2025

LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs

Leqi Shen, Tao He, Guoqiang Gong +5

Training-free video large language models (LLMs) leverage pretrained Image LLMs to process video content without the need for further training. A key challenge in such approaches i…