3 papers
cs.LG2026
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
Fang Wu, Aaron Tu, Weihao Xuan +21
Reinforcement learning with verifiable rewards (RLVR) is a practical, scalable way to improve large language models on math, code, and other structured tasks. However, we argue tha…
cs.AI2025
When Reasoning Meets Its Laws
Junyu Zhang, Yifan Sun, Tianang Leng +4
Despite the superior performance of Large Reasoning Models (LRMs), their reasoning behaviors are often counterintuitive, leading to suboptimal reasoning capabilities. To theoretica…
cs.LG2025
Predicting and generating antibiotics against future pathogens with ApexOracle
Tianang Leng, Fangping Wan, Marcelo Der Torossian Torres +1
Antimicrobial resistance (AMR) is escalating and outpacing current antibiotic development. Thus, discovering antibiotics effective against emerging pathogens is becoming increasing…