Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
MoE-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation
Qingyu Yang, Haonan He, Minglei Li +4
Mixture-of-Experts (MoE) architectures have been widely adopted in large language models, yet parameter-efficient fine-tuning (PEFT) for MoE models remains underexplored. Existing…
cs.CL2026
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
Lei Bai, Zongsheng Cao, Yang Chen +50
We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We investigate agent-horizon scaling…
cs.CL2026
Parametric Skills
Xuan Zhao, Haonan He, Qingyu Yang +5
Since intelligence fundamentally relies on efficient skill acquisition (Chollet, 2019), the ability to leverage skills is critical. For LLMs, skills, manually authored or extracted…