2 papers
cs.LG2026
FPGA Co-Design for Efficient N:M Sparse and Quantized Model Inference
Fen-Yu Hsieh, Yun-Chang Teng, Ding-Yong Hong +1
Large language models (LLMs) have demonstrated remarkable performance across a wide range of language processing tasks. However, this success comes at the cost of substantial compu…
cs.CL2025
SAP: Syntactic Attention Pruning for Transformer-based Language Models
Tzu-Yun Lee, Ding-Yong Hong, Jan-Jan Wu
This paper introduces Syntactic Attention Pruning (SAP), a novel method for effectively pruning attention heads in Transformer models. Unlike conventional approaches that rely sole…