Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
Minbin Huang, Han Shi, Chuanyang Zheng +5
Modern Mixture-of-Experts (MoE) architectures allocate expert capacity through a rigid per-layer rule: each transformer layer owns a separate expert set. This convention couples de…
cs.LG2025
Logits-Based Finetuning
Jingyao Li, Senqiao Yang, Sitong Wu +4
In recent years, developing compact and efficient large language models (LLMs) has emerged as a thriving area of research. Traditional Supervised Fine-Tuning (SFT), which relies on…