2 papers
cs.CL2026
Benchmark Illusion: Disagreement among LLMs and Its Scientific Consequences
Eddie Yang, Dashun Wang
Benchmarks underpin how progress in large language models (LLMs) is measured and trusted. Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep e…
physics.soc-ph2026
Pivoting as an Adaptive Strategy to Geopolitical Tensions in U.S. Science
Moxin Li, Yifang Ma, Yang Wang +1
Geopolitical tensions increasingly reshape the structure and openness of global science, yet we still lack a clear understanding of how successfully scientists adapt their work und…