4 papers
Differentiable Particle-Mesh Ewald with Cartesian Tensor Message Passing for Learning Long-Range Electrostatics and Dipole Response
Zhiyue Guo, Junjie Wang, Haoting Zhang +4
Machine learning interatomic potentials (MLIPs) can approach quantum accuracy for short-range chemistry, but most architectures remain local and fail to capture the long-range elec…
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
Yijie Xu, Huizai Yao, Zhiyu Guo +5
Large language models (LLMs) are increasingly deployed in specialized domains such as finance, medicine, and agriculture, where they face significant distribution shifts from their…
Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters
Zhiyu Guo, Hidetaka Kamigaito, Taro Watanabe
Scaling the context size of large language models (LLMs) enables them to perform various new tasks, e.g., book summarization. However, the memory cost of the Key and Value (KV) cac…
Dependency-Aware Semi-Structured Sparsity of GLU Variants in Large Language Models
Zhiyu Guo, Hidetaka Kamigaito, Taro Wanatnabe
The rapid advancement in Large Language Models (LLMs) has markedly enhanced the capabilities of language understanding and generation. However, the substantial model size poses har…