5 papers
Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic Sampling
Siyuan Gan, Yuhan Li, Xiran Wang +6
Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) is a prominent variant of Group Relative Policy Optimization (GRPO). DAPO introduces several improvements over GRPO.…
When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation
Siyuan Gan, Yuhan Li, Xiran Wang +5
On-policy distillation (OPD) has recently emerged as a popular post-training paradigm for large language models (LLMs), providing an efficient way to transfer the knowledge and cap…
Exploring Dualistic Meta-Learning to Enhance Domain Generalization in Open Set Scenarios
Xiran Wang, Jian Zhang, Lei Qi +2
Domain generalization learns from multiple source domains to generalize to unseen target domains. However, it often neglects the realistic case of label mismatch between source and…
JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency
Aichen Cai, Anmeng Zhang, Anyu Li +66
We introduce JoyAI-LLM Flash, an efficient Mixture-of-Experts (MoE) language model designed to redefine the trade-off between strong performance and token efficiency in the sub-50B…
Balanced Direction from Multifarious Choices: Arithmetic Meta-Learning for Domain Generalization
Xiran Wang, Jian Zhang, Lei Qi +1
Domain generalization is proposed to address distribution shift, arising from statistical disparities between training source and unseen target domains. The widely used first-order…