4 papers · 1 filter
BRIDGE: Predicting Human Task Completion Time From Model Performance
Fengyuan Liu, Jay Gala, Nilaksh +3
Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty. Existing approaches that rely on d…
The Alien Space of Science: Sampling Coherent but Cognitively Unavailable Research Directions
Alejandro H. Artiles, Martin Weiss, Levin Brinkmann +6
Scientific discovery is constrained not only by what is true, but by what is cognitively available to the researchers currently exploring a field. Many directions are coherent in l…
Capturing Individual Human Preferences with Reward Features
André Barreto, Vincent Dumoulin, Yiran Mao +6
Reinforcement learning from human feedback usually models preferences using a reward function that does not distinguish between people. We argue that this is unlikely to be a good…
ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review
Gaurav Sahu, Hugo Larochelle, Laurent Charlin +1
Peer review is the cornerstone of scientific publishing, yet it suffers from inconsistencies, reviewer subjectivity, and scalability challenges. We introduce ReviewerToo, a modular…