4 papers
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Contemporary AI lacks the imagination to diverge or negate in science
Honglin Bao, Siyang Wu, Xiao Liu +3
Bold claims that AI will accelerate scientific discovery have raced ahead of evidence from working scientists, yet large-scale, scientist-in-the-loop evidence is scarce. Here we mo…
Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction
Zehan Li, Yutong Zhu, Siyang Wu +2
Large language models produce fluent fiction, yet their creative output is widely seen as flat. We ask where this quality originates in the training and whether it affects differen…
Building an Atlas of Social Experiments to Link Studies, Reconcile Conflicts, and Bridge Gaps
Jiawei Zhang, Honglin Bao, Pengda Wang +3
Social and behavioral science runs thousands of experiments each year, yet their findings rarely accumulate into a coherent map of what is known, what conflicts, and what remains m…