3 papers
cs.LG2026
SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter
Powei Chang, Jinpeng Zhang, Chaoqun Sun +6
Reinforcement learning with verifiable rewards (RLVR) often adopts GRPO-style group-relative updates, sampling multiple rollouts per prompt to construct normalized learning signals…
cs.LG2026
ChemAmp: Amplified Chemistry Tools via Composable Agents
Zhucong Li, Powei Chang, Jin Xiao +6
Although LLM-based agents are proven to master tool orchestration in scientific fields, particularly chemistry, their single-task performance remains limited by underlying tool con…
cs.LG2026
SPICE: Submodular Penalized Information-Conflict Selection for Efficient Large Language Model Training
Powei Chang, Jinpeng Zhang, Bowen Chen +9
Information-based data selection for instruction tuning is compelling: maximizing the log-determinant of the Fisher information yields a monotone submodular objective, enabling gre…