2 papers
cs.LG2026
Data-Efficient RLVR via Off-Policy Influence Guidance
Erle Zhu, Dazhi Jiang, Yuan Wang +8
Data selection is a critical aspect of Reinforcement Learning with Verifiable Rewards (RLVR) for enhancing the reasoning capabilities of large language models (LLMs). Current data…
cs.AI2024
Configurable Foundation Models: Building LLMs from a Modular Perspective
Chaojun Xiao, Zhengyan Zhang, Chenyang Song +20
Advancements in LLMs have recently unveiled challenges tied to computational efficiency and continual scalability due to their requirements of huge parameters, making the applicati…