5 papers
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
Pujun Zheng, Zixin Shang, Shufan Jiang +5
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluat…
GraphReview: Scientific Paper Evaluation via LLM-based Graph Evidence Expansion
Pujun Zheng, Wanying Ren, Jiacheng Yao +2
Scientific paper evaluation often involves not only assessing a manuscript itself, but also relating it to contemporaneous research and prior literature. However, existing LLM-base…
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
Jinquan Zheng, Jia Yuan, Jiacheng Yao +3
Large language models (LLMs) used for multiple-choice and pairwise evaluation tasks often exhibit selection bias due to non-semantic factors like option positions and label symbols…
MoRI: Learning Motivation-Grounded Reasoning for Scientific Ideation in Large Language Models
Chenyang Gu, Jiahao Cheng, Meicong Zhang +3
Scientific ideation aims to propose novel solutions within a given scientific context. Existing LLM-based agentic approaches emulate human research workflows, yet inadequately mode…
From Isolated Scoring to Collaborative Ranking: A Comparison-Native Framework for LLM-Based Paper Evaluation
Pujun Zheng, Jiacheng Yao, Jinquan Zheng +6
Large language models (LLMs) are currently applied to scientific paper evaluation by assigning an absolute score to each paper independently. However, since score scales vary acros…