2 citations · 3 across the 8 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling
Runsong Zhao, Shilei Liu, Jiwei Tang +8
The quadratic complexity and indefinitely growing key-value (KV) cache of standard Transformers pose a major barrier to long-context processing. To overcome this, we introduce the…
cs.LG2026
Expert Divergence Learning for MoE-based Language Models
Jiaang Li, Haibin Chen, Langming Liu +9
The Mixture-of-Experts (MoE) architecture is a powerful technique for scaling language models, yet it often suffers from expert homogenization, where experts learn redundant functi…