collaborators

7 papers

cs.CV2026

Adaptive MLP Pruning for Large Vision Transformers

Chengchao Shen

Large vision transformers present impressive scalability, as their performance can be well improved with increased model capacity. Nevertheless, their cumbersome parameters results…

cs.CL2026

High-Fidelity Pruning for Large Language Models

Yijun Zhu, Jianxin Wang, Chengchao Shen

Large Language Models (LLMs) have demonstrated exceptional performance across a wide range of tasks, yet their significant computational and memory requirements present major chall…

cs.CV2026

Hallucination Begins Where Saliency Drops

Xiaofeng Zhang, Yuanchao Zhu, Chaochen Gu +8

Recent studies have examined attention dynamics in large vision-language models (LVLMs) to detect hallucinations. However, existing approaches remain limited in reliably distinguis…

cs.CV2025

Diversity-Guided MLP Reduction for Efficient Large Vision Transformers

Chengchao Shen, Hourun Zhu, Gongfan Fang +2

Transformer models achieve excellent scaling property, where the performance is improved with the increment of model capacity. However, large-scale model parameters lead to an unaf…

cs.LG2025

Optimal Corpus Aware Training for Neural Machine Translation

Yi-Hsiu Liao, Cheng Shen, Brenda +1

Corpus Aware Training (CAT) leverages valuable corpus metadata during training by injecting corpus information into each training example, and has been found effective in the liter…

cs.CL2025

SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models

Hourun Zhu, Chengchao Shen

In spite of strong performance achieved by LLMs, the costs of their deployment are unaffordable. For the compression of LLMs, gradient-based pruning methods present promising effec…