1 citations · 1 across the 5 of their papers we have counts for
5 papers
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
Jiayu Qin, Jianchao Tan, Kefeng Zhang +2
The remarkable performance of large language models (LLMs) in various language tasks has attracted considerable attention. However, the ever-increasing size of these models present…
C2T: A Classifier-Based Tree Construction Method in Speculative Decoding
Feiye Huo, Jianchao Tan, Kefeng Zhang +2
The growing scale of Large Language Models (LLMs) has exacerbated inference latency and computational costs. Speculative decoding methods, which aim to mitigate these issues, often…
ASP: Automatic Selection of Proxy dataset for efficient AutoML
Peng Yao, Chao Liao, Jiyuan Jia +4
Deep neural networks have gained great success due to the increasing amounts of data, and diverse effective neural network designs. However, it also brings a heavy computing burden…
USDC: Unified Static and Dynamic Compression for Visual Transformer
Huan Yuan, Chao Liao, Jianchao Tan +5
Visual Transformers have achieved great success in almost all vision tasks, such as classification, detection, and so on. However, the model complexity and the inference speed of t…
Resource Constrained Model Compression via Minimax Optimization for Spiking Neural Networks
Jue Chen, Huan Yuan, Jianchao Tan +3
Brain-inspired Spiking Neural Networks (SNNs) have the characteristics of event-driven and high energy-efficient, which are different from traditional Artificial Neural Networks (A…