39 citations · 53 across the 17 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Data-Efficient RLVR via Off-Policy Influence Guidance
Erle Zhu, Dazhi Jiang, Yuan Wang +8
Data selection is a critical aspect of Reinforcement Learning with Verifiable Rewards (RLVR) for enhancing the reasoning capabilities of large language models (LLMs). Current data…
cs.LG2021★ 39 cited
FastMoE: A Fast Mixture-of-Expert Training System
Jiaao He, Jiezhong Qiu, Aohan Zeng +3
Mixture-of-Expert (MoE) presents a strong potential in enlarging the size of language model to trillions of parameters. However, training trillion-scale MoE requires algorithm and…