3 papers
cs.LG2026
Metag: A dataset to build agentic meta-reviewing capabilities
Anirudh Sundar, Min Chen, Divya Tadimeti +11
AI tools increasingly support tasks across the scientific research cycle, from experiment design and manuscript preparation to peer review. At the same time, the continuing growth…
cs.SE2025
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
Ananya Singha, Harshita Sahijwani, Walt Williams +10
Excel is a pervasive yet often complex tool, particularly for novice users, where runtime errors arising from logical mistakes or misinterpretations of functions pose a significant…
cs.CL2024
SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text
Reshmi Ghosh, Tianyi Yao, Lizzy Chen +5
Large Language Model (LLM) integrations into applications like Microsoft365 suite and Google Workspace for creating/processing documents, emails, presentations, etc. has led to con…