3 papers
cs.AI2026
Generic Expert Coverage for Pruning SparseMixture-of-Experts Language Models
Yongqin Zeng, Sicheng Pan, Jiale Wang +4
Sparsely activated Mixture-of-Experts (MoE) language models contain substantial structured redundancy among routed experts, but pruning them without downstream calibration data rem…
cs.CL2026
CoreMem: Riemannian Retrieval and Fisher-Guided Distillation for Long-Term Memory in Dialogue Agents
Jiaqi Chen, Yongqin Zeng, Shaoshen Chen +4
Personalized dialogue agents require continuous long-term memory to maintain coherent interactions across multiple sessions. However, deploying these capabilities on consumer-grade…
cs.CL2026
Experience-Driven Dynamic Exits for LLMs with Reinforcement Learning
Yanyu Zhu, Hoilam Pao, Niu Hu +6
Large Language Models suffer from slow autoregressive inference. While self-speculative decoding accelerates this process, its efficiency is hampered by static configurations like…