17 papers
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
Fengqing Jiang, Yichen Feng, Yuetai Li +3
The convergence of LLM-powered research assistants and AI-based peer review systems creates a critical vulnerability: fully automated publication loops where AI-generated research…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations
Penghui Yang, Zhonghan Zhang, Yue Li +6
Density functional theory (DFT) serves as the basis for computational discovery in materials science and chemistry, yet each calculation demands extensive human effort: adjusting a…
JobBench: Aligning Agent Work With Human Will
Yuetai Li, Yichen Feng, Zhangchen Xu +21
Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluates AI agents on the workflow…
Polyhedral Instability Governs Regret in Online Learning
Yuetai Li, Fengqing Jiang, Yichen Feng +6
Many online decision problems over combinatorial actions are addressed via convex relaxations, leading to online convex optimization with piecewise linear objectives and induced po…
The WidthWall: A Strict Expressivity Hierarchy for Hypergraph Neural Networks
Fengqing Jiang, Yuetai Li, Yichen Feng +6
Hypergraphs provide a natural framework to model higher-order interactions in scientific, social, and biological systems. Hypergraph neural networks (HGNNs) aim to learn from such…