2 papers
cs.CL2026
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
Bingyang Ye, Shan Chen, Jingxuan Tu +4
Large language models are increasingly being used to assess and forecast research ideas, yet we lack scalable ways to evaluate the quality of models' judgments about these scientif…
cs.AI2025
Large language models require a new form of oversight: capability-based monitoring
Katherine C. Kellogg, Bingyang Ye, Yifan Hu +3
The rapid adoption of large language models (LLMs) in healthcare has been accompanied by scrutiny of their oversight. Existing monitoring approaches, inherited from traditional mac…