4 papers
Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data
Renfei Zhang, Niloofar Mireshghallah
Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what models will divulge is under studied. Here w…
SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?
Kevin Han, Renfei Zhang, Kathy Wei +3
LLM agents have incredible potential for scientific discovery applications. However, the performance of LLM agents on real-world, small molecule drug design (SMDD) tasks across div…
Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs
Renfei Zhang, Manasa Kaniselvan, Rylan Schaeffer +1
Reinforcement learning (RL) is often credited with improving reasoning at the expense of factual knowledge. We instead find that reasoning models outperform their instruction-tuned…
Why Pool When You Can Flow? Active Learning with GFlowNets
Renfei Zhang, Mohit Pandey, Artem Cherkasov +1
The scalability of pool-based active learning is limited by the computational cost of evaluating large unlabeled datasets, a challenge that is particularly acute in virtual screeni…