5 papers
AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery
Yu Li, Chenyang Shao, Xinyang Liu +13
Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-The-Art (SOTA) performance, creat…
GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces
Xinyu Geng, Yanjing Xiao, Yuyang Zhang +5
Deep research agents integrate fragmented evidence through multi-step tool use. BrowseComp offers a text-only testbed for such agents, but existing multimodal benchmarks rarely req…
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
Xuan Liu, Haoyang Shang, Zizhang Liu +4
Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable and highly sensitive to design choi…
OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
Chenyang Shao, Dehao Huang, Yu Li +18
With the rapid development of Large Language Models (LLMs), AI agents have demonstrated increasing proficiency in scientific tasks, ranging from hypothesis generation and experimen…
Route-and-Reason: Scaling Large Language Model Reasoning with Reinforced Model Router
Chenyang Shao, Xinyang Liu, Yutang Lin +2
Chain-of-thought has been proven essential for enhancing the complex reasoning abilities of Large Language Models (LLMs), but it also leads to high computational costs. Recent adva…