2 papers
cs.CL2025
Cross-layer Attention Sharing for Pre-trained Large Language Models
Yongyu Mu, Yuzhang Wu, Yuchun Fan +9
To enhance the efficiency of the attention mechanism within large language models (LLMs), previous works primarily compress the KV cache or group attention heads, while largely ove…
cs.CL2025
Revealing the Parallel Multilingual Learning within Large Language Models
Yongyu Mu, Peinan Feng, Zhiquan Cao +8
In this study, we reveal an in-context learning (ICL) capability of multilingual large language models (LLMs): by translating the input to several languages, we provide Parallel In…