3 papers
cs.CL2025
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
Siyuan Yan, Guo-Qing Jiang, Yuchen Zhang +4
Large language models (LLMs) now support context windows of hundreds of thousands to millions of tokens, enabling applications such as long-document summarization, large-scale code…
cs.LG2025
Exploiting Block Coordinate Descent for Cost-Effective LLM Model Training
Zeyu Liu, Yan Li, Yunquan Zhang +6
Training large language models typically demands extensive GPU memory and substantial financial investment, which poses a barrier for many small- to medium-sized teams. In this pap…
cs.CL2025
Scaling Laws for Speculative Decoding
Siyuan Yan, Mo Zhu, Guo-qing Jiang +8
The escalating demand for efficient decoding in large language models (LLMs) is particularly critical for reasoning-intensive architectures like OpenAI-o3 and DeepSeek-R1, which de…