2 papers
cs.LG2026
Convex Dataset Valuation for Post-Training
Siqi Zeng, Christopher Jung, Rui Li +7
Improving LLM performance on downstream tasks sometimes requires leveraging auxiliary datasets during post-training. In practice, however, developers face constraints on compute, l…
cs.CL2025
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
Linglin Jing, Yuting Gao, Zhigang Wang +5
Recent advancements have shown that the Mixture of Experts (MoE) approach significantly enhances the capacity of large language models (LLMs) and improves performance on downstream…