Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
SAP: Syntactic Attention Pruning for Transformer-based Language Models
Tzu-Yun Lee, Ding-Yong Hong, Jan-Jan Wu
This paper introduces Syntactic Attention Pruning (SAP), a novel method for effectively pruning attention heads in Transformer models. Unlike conventional approaches that rely sole…
cs.CL2025
AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference
Kuan-Wei Lu, Ding-Yong Hong, Pangfeng Liu +1
Large language models (LLMs) have achieved remarkable performance across a wide range of tasks, but their increasing parameter sizes significantly slow down inference. Speculative…