2 papers
cs.LG2026
LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining
Qiuwu Chen, Zimo Liu, Yuchen Li +8
Large language models (LLMs) have achieved remarkable breakthroughs across various applications. However, their architectures remain inefficient in pretraining due to two main limi…
cs.AI2026
cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs
Xin Yang, Yemin Wang, Mingda Liu +4
Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training…