2 papers
cs.CL2025
FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
Sarina Xi, Vishisht Rao, Justin Payan +1
The identification and localization of errors is a core task in peer review, yet the exponential growth of scientific output has made it increasingly difficult for human reviewers…
cs.HC2025
Who is a Better Matchmaker? Human vs. Algorithmic Judge Assignment in a High-Stakes Startup Competition
Sarina Xi, Orelia Pi, Miaomiao Zhang +3
There is growing interest in applying artificial intelligence (AI) to automate and support complex decision-making tasks. However, it remains unclear how algorithms compare to huma…