2 papers
cs.CL2025
Consultant Decoding: Yet Another Synergistic Mechanism
Chuanghao Ding, Jiaping Wang, Ziqing Yang +4
The synergistic mechanism based on Speculative Decoding (SD) has garnered considerable attention as a simple yet effective approach for accelerating the inference of large language…
cs.LG2025
LeetDecoding: A PyTorch Library for Exponentially Decaying Causal Linear Attention with CUDA Implementations
Jiaping Wang, Simiao Zhang, Qiao-Chu He +1
The machine learning and data science community has made significant while dispersive progress in accelerating transformer-based large language models (LLMs), and one promising app…