2 papers
cs.CL2026
Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs
Rongfeng Wang, Shichao Weng, Zhiqiang Wang +4
Mixture-of-Experts (MoE) models route each token to a subset of expert networks, increasing capacity while keeping per-token computation sparse. In many deployed MoEs, the number o…
cs.SE2026
Learning Globally Reusable Skills for Coding Agents
Chen Yang, Jiashuo Tian, Ziqi Wang +3
Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing approaches typically treat skill evoluti…