5 papers
Bayesian and Motivated Reasoning in AI Agents
Eddie Yang
AI agents increasingly perform open-ended tasks in settings where their conclusions can guide consequential decisions. We provide evidence that AI agents draw different conclusions…
BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery
Jieyi Wang, Bingxuan Li, Nanyi Jiang +9
Biomedical deep-research systems increasingly retrieve and synthesize scientific evidence, but their outputs typically collapse heterogeneous evidence into static text, making prov…
Sci2Pol: Evaluating and Fine-tuning LLMs on Scientific-to-Policy Brief Generation
Weimin Wu, Alexander C. Furnas, Eddie Yang +5
We propose Sci2Pol-Bench and Sci2Pol-Corpus, the first benchmark and training dataset for evaluating and fine-tuning large language models (LLMs) on policy brief generation from a…
Benchmark Illusion: Disagreement among LLMs and Its Scientific Consequences
Eddie Yang, Dashun Wang
Benchmarks underpin how progress in large language models (LLMs) is measured and trusted. Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep e…
Discovering influential text using convolutional neural networks
Megan Ayers, Luke Sanford, Margaret Roberts +1
Experimental methods for estimating the impacts of text on human evaluation have been widely used in the social sciences. However, researchers in experimental settings are usually…