Had enough of experts? Quantitative knowledge retrieval from large language models
arXiv:2402.07770 · doi:10.1002/sta4.70054
Abstract
Large language models (LLMs) have been extensively studied for their abilities to generate convincing natural language sequences, however their utility for quantitative information retrieval is less well understood. Here we explore the feasibility of LLMs as a mechanism for quantitative knowledge retrieval to aid two data analysis tasks: elicitation of prior distributions for Bayesian models and imputation of missing data. We introduce a framework that leverages LLMs to enhance Bayesian workflows by eliciting expert-like prior knowledge and imputing missing data. Tested on diverse datasets, this approach can improve predictive accuracy and reduce data requirements, offering significant potential in healthcare, environmental science and engineering applications. We discuss the implications and challenges of treating LLMs as 'experts'.
References in corpus (20)
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- The prior can generally only be understood in the context of the likelihood
- How Generative AI models such as ChatGPT can be (Mis)Used in SPC Practice, Education, and Research? An Exploratory Study
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
- OpenAGI: When LLM Meets Domain Experts
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- HAGRID: A Human-LLM Collaborative Dataset for Generative Information-Seeking with Attribution
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
- Solving Math Word Problems by Combining Language Models With Symbolic Solvers
- Numeracy from Literacy: Data Science as an Emergent Skill from Large Language Models
- Large Language Models as Data Preprocessors
- Eliciting Human Preferences with Language Models
- Table-GPT: Table-tuned GPT for Diverse Table Tasks
- GATGPT: A Pre-trained Large Language Model with Graph Attention Network for Spatiotemporal Imputation
- Task Contamination: Language Models May Not Be Few-Shot Anymore
- Jellyfish: A Large Language Model for Data Preprocessing
- RetClean: Retrieval-Based Data Cleaning Using Foundation Models and Data Lakes
- Foundations of Bayesian Learning from Synthetic Data
- A Context-Aware Approach for Enhancing Data Imputation with Pre-trained Language Models
- AutoElicit: Using Large Language Models for Expert Prior Elicitation in Predictive Modelling