2 papers
cs.CL2026
LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks
Xiao Zhang, Qumeng Sun, Jiahao Li +4
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy st…
cs.AI2026
AutoResearch: Insight In, Hallucination Out
Yiming Ren, Xiang Liu, Qumeng Sun +4
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically gr…