Large language models surpass human experts in predicting neuroscience results
arXiv:2403.03230 · doi:10.1038/s41562-024-02046-9
Abstract
Scientific discoveries often hinge on synthesizing decades of research, a task that potentially outstrips human information processing capacities. Large language models (LLMs) offer a solution. LLMs trained on the vast scientific literature could potentially integrate noisy yet interrelated findings to forecast novel results better than human experts. To evaluate this possibility, we created BrainBench, a forward-looking benchmark for predicting neuroscience results. We find that LLMs surpass experts in predicting experimental outcomes. BrainGPT, an LLM we tuned on the neuroscience literature, performed better yet. Like human experts, when LLMs were confident in their predictions, they were more likely to be correct, which presages a future where humans and LLMs team together to make discoveries. Our approach is not neuroscience-specific and is transferable to other knowledge-intensive endeavors.
The latest version of this paper has been published at Nature Human Behaviour, please see https://www.nature.com/articles/s41562-024-02046-9
References in corpus (8)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Summary of ChatGPT-Related Research and Perspective Towards the Future of Large Language Models
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Measuring Massive Multitask Language Understanding
- Extracting Training Data from Large Language Models
- Scalable Extraction of Training Data from (Production) Language Models