2 papers
cs.LG2026
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
Wenzheng Zhang, Bingzheng Liu, Yang Hu +3
Quantization-Aware Training from scratch has emerged as a promising approach for building efficient large language models (LLMs) with extremely low-bit weights (sub 2-bit), which c…
cs.LG2025
Efficient Linear Attention for Multivariate Time Series Modeling via Entropy Equality
Mingtao Zhang, Guoli Yang, Zhanxing Zhu +2
Attention mechanisms have been extensively employed in various applications, including time series modeling, owing to their capacity to capture intricate dependencies; however, the…