3 papers
cs.CL2025
Entropy-Guided Reasoning Compression
Hourun Zhu, Yang Gao, Wenlong Fei +2
Large reasoning models have demonstrated remarkable performance on complex reasoning tasks, yet the excessive length of their chain-of-thought outputs remains a major practical bot…
cs.CV2025
Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
Chengchao Shen, Hourun Zhu, Gongfan Fang +2
Transformer models achieve excellent scaling property, where the performance is improved with the increment of model capacity. However, large-scale model parameters lead to an unaf…
cs.CL2025
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
Hourun Zhu, Chengchao Shen
In spite of strong performance achieved by LLMs, the costs of their deployment are unaffordable. For the compression of LLMs, gradient-based pruning methods present promising effec…