2 citations · 2 across the 5 of their papers we have counts for
4 papers · 1 filter
DE-Venus: A Data-Efficient RLVR Framework for Large Language Models
Shenzhi Yang, Guangcheng Zhu, Kai Tang +11
Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost…
REAR: Test-time Preference Realignment through Reward Decomposition
Fuxiang Zhang, Pengcheng Wang, Chenran Li +6
Aligning large language models (LLMs) with diverse user preferences is a critical yet challenging task. While post-training methods can adapt models to specific needs, they often r…
AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations
Penghui Yang, Zhonghan Zhang, Yue Li +6
Density functional theory (DFT) serves as the basis for computational discovery in materials science and chemistry, yet each calculation demands extensive human effort: adjusting a…
Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning
Xin Cheng, Shuo He, Lang Feng +4
Group-based reinforcement learning (RL) methods have achieved remarkable success in improving the performance of large language models (LLMs) and have been rapidly extended to agen…