2 papers
cs.AI2026
Generic Expert Coverage for Pruning SparseMixture-of-Experts Language Models
Yongqin Zeng, Sicheng Pan, Jiale Wang +4
Sparsely activated Mixture-of-Experts (MoE) language models contain substantial structured redundancy among routed experts, but pruning them without downstream calibration data rem…
cs.CL2026
CoreMem: Riemannian Retrieval and Fisher-Guided Distillation for Long-Term Memory in Dialogue Agents
Jiaqi Chen, Yongqin Zeng, Shaoshen Chen +4
Personalized dialogue agents require continuous long-term memory to maintain coherent interactions across multiple sessions. However, deploying these capabilities on consumer-grade…