10 papers
BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution
Yangzhen Wu, Aaron J. Li, Wenjie Ma +10
The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differentiate model capabilities or pr…
Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
Zhou Ziheng, Huacong Tang, Jinyuan Zhang +9
Discovering causal regularities and applying them to build functional systems--the discovery-to-application loop--is a hallmark of general intelligence, yet evaluating this capacit…
Why Are We Moral? An LLM-based Agent Simulation Approach to Study Moral Evolution
Zhou Ziheng, Huacong Tang, Mingjie Bi +7
The evolution of morality presents a puzzle: natural selection should favor self-interest, yet humans developed moral systems promoting altruism. Traditional approaches must abstra…
PRBench: End-to-end Paper Reproduction in Physics Research
Shi Qiu, Junyi Deng, Yiwei Deng +48
AI agents powered by large language models exhibit strong reasoning and problem-solving capabilities, enabling them to assist scientific research tasks such as formula derivation a…
How do Role Models Shape Collective Morality? Exemplar-Driven Moral Learning in Multi-Agent Simulation
Junjie Liao, Huacong Tang, Zhou Ziheng +2
Do We Need Role Models? How do Role Models Shape Collective Morality? To explore the questions, we build a multi-agent simulation powered by a Large Language Model, where agents wi…
Credibility Governance: A Social Mechanism for Collective Self-Correction under Weak Truth Signals
Wanying He, Yanxi Lin, Ziheng Zhou +5
Online platforms increasingly rely on opinion aggregation to allocate real-world attention and resources, yet common signals such as engagement votes or capital-weighted commitment…