2 papers
cs.CL2026
Humans and LLMs Diverge on Probabilistic Inferences
Gaurav Kamath, Sreenath Madathil, Sebastian Schuster +2
Human reasoning often involves working over limited information to arrive at probabilistic conclusions. In its simplest form, this involves making an inference that is not strictly…
cs.CL2025
RExBench: Can coding agents autonomously implement AI research extensions?
Nicholas Edwards, Yukyung Lee, Yujun Audrey Mao +3
Agents based on Large Language Models (LLMs) have shown promise for performing sophisticated software engineering tasks autonomously. In addition, there has been progress towards d…