11 papers
PreScience: A Dataset and Benchmark for Scientific Forecasting
Anirudh Ajith, Amanpreet Singh, Jay DeYoung +7
Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark for scientific forecasting built a…
Measuring Behavior Portability in Large Language Models
Tianjia Dong, Nadav Kunievsky, James A. Evans
Large language models are increasingly deployed as autonomous decision makers, yet the behavioral mapping they exhibit can vary substantially across decision environments that are…
Measuring Intent Comprehension in LLMs
Nadav Kunievsky, James A. Evans
People judge interactions with large language models (LLMs) as successful when outputs match what they want, not what they type. Yet LLMs are trained to predict the next token sole…
U.S. Technological Containment and the Rise of China's Open AI Ecosystem
Wang Jin, Nadav Kunievsky, Bowen Lou +2
Over the past decade, U.S. policies have increasingly aimed to preserve artificial intelligence (AI) leadership by promoting domestic free-market policies while controlling global…
The Effect of Age at Arrival on the Alignment Between Immigrant and Native-Born Gender Norms: A Distributional Approach
Nadav Kunievsky
This paper examines how age at migration affects cultural assimilation by studying convergence in gender role attitudes between immigrants and the UK-born population. Although cult…
Missing vs. Unused Knowledge Hypothesis for Language Model Bottlenecks in Patent Understanding
Siyang Wu, Honglin Bao, Nadav Kunievsky +1
While large language models (LLMs) excel at factual recall, the real challenge lies in knowledge application. A gap persists between their ability to answer complex questions and t…