3 papers
cs.AI2026
Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment
Zhiyu Chen, Keyu Zhao, Jigao Fu +8
Evaluating research ideas generated by LLMs is difficult because their scientific value cannot be fully determined by objective criteria, and no single reference answer specifies w…
cs.AI2026
SETA: Scaling Environments for Terminal Agents
Qijia Shen, Zhiqi Huang, Vamsidhar Kamanuru +19
Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical user interfaces (GUIs). Among these, the te…
cs.AI2026
LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform
Ruotong Zhao, Zhiyu Chen, Xurui Liu +7
Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend…