Validation Requirements for AI-based Intervention-Evaluation in Aging and Longevity Research and Practice
arXiv:2408.15264 · doi:10.1016/j.arr.2024.102617
Abstract
The field of aging and longevity research is overwhelmed by vast amounts of data, calling for the use of Artificial Intelligence (AI), including Large Language Models (LLMs), for the evaluation of geroprotective interventions. Such evaluations should be correct, useful, comprehensive, explainable, and they should consider causality, interdisciplinarity, adherence to standards, longitudinal data and known aging biology. In particular, comprehensive analyses should go beyond comparing data based on canonical biomedical databases, suggesting the use of AI to interpret changes in biomarkers and outcomes. Our requirements motivate the use of LLMs with Knowledge Graphs and dedicated workflows employing, e.g., Retrieval-Augmented Generation. While naive trust in the responses of AI tools can cause harm, adding our requirements to LLM queries can improve response quality, calling for benchmarking efforts and justifying the informed use of LLMs for advice on longevity interventions.
11 pages, 1 Figure, 1 Table
References in corpus (11)
- Unifying Large Language Models and Knowledge Graphs: A Roadmap
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Large Language Models Streamline Automated Machine Learning for Clinical Studies
- Gorilla: Large Language Model Connected with Massive APIs
- Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity
- Plex: Towards Reliability using Pretrained Large Model Extensions
- Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering
- Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
- Can LLMs Produce Faithful Explanations For Fact-checking? Towards Faithful Explainable Fact-Checking via Multi-Agent Debate
- Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models
- Learning to Check: Unleashing Potentials for Self-Correction in Large Language Models