Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
Sarina Xi, Vishisht Rao, Justin Payan +1
The identification and localization of errors is a core task in peer review, yet the exponential growth of scientific output has made it increasingly difficult for human reviewers…
cs.CL2023
InstructExcel: A Benchmark for Natural Language Instruction in Excel
Justin Payan, Swaroop Mishra, Mukul Singh +7
With the evolution of Large Language Models (LLMs) we can solve increasingly more complex NLP tasks across various domains, including spreadsheets. This work investigates whether L…