2 papers
cs.CL2026
AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research
Talor Abramovich, Gal Chechik
Language model agents are increasingly used to automate scientific research, yet evaluating their scientific contributions remains a challenge. A key mechanism to obtain such insig…
cs.LG2025
Policy Gradient with Tree Expansion
Gal Dalal, Assaf Hallak, Gugan Thoppe +2
Policy gradient methods are notorious for having a large variance and high sample complexity. To mitigate this, we introduce SoftTreeMax -- a generalization of softmax that employs…