1 citations · 1 across the 9 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
Ning Yang, Hengyu Zhong, Wentao Wang +3
The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by lightweight Continual Pre-Training (CPT). While effec…
cs.CL2026
Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression
Peijun Zhu, Ning Yang, Baoliang Tian +4
Mixture-of-Experts (MoE) Large Language Models (LLMs) face a trilemma of load imbalance, parameter redundancy, and communication overhead. We introduce a unified framework based on…