A Reliable Knowledge Processing Framework for Combustion Science using Foundation Models
arXiv:2401.00544 · doi:10.1016/j.egyai.2024.100365
Abstract
This research explores the integration of large language models (LLMs) into scientific data assimilation, focusing on combustion science as a case study. Leveraging foundational models integrated with Retrieval-Augmented Generation (RAG) framework, the study introduces an approach to process diverse combustion research data, spanning experimental studies, simulations, and literature. The multifaceted nature of combustion research emphasizes the critical role of knowledge processing in navigating and extracting valuable information from a vast and diverse pool of sources. The developed approach minimizes computational and economic expenses while optimizing data privacy and accuracy. It incorporates prompt engineering and offline open-source LLMs, offering user autonomy in selecting base models. The study provides a thorough examination of text segmentation strategies, conducts comparative studies between LLMs, and explores various optimized prompts to demonstrate the effectiveness of the framework. By incorporating an external database, the framework outperforms a conventional LLM in generating accurate responses and constructing robust arguments. Additionally, the study delves into the investigation of optimized prompt templates for the purpose of efficient extraction of scientific literature. The research addresses concerns related to hallucinations and false research articles by introducing a custom workflow developed with a detection algorithm to filter out inaccuracies. Despite identified areas for improvement, the framework consistently delivers accurate domain-specific responses with minimal human oversight. The prompt-agnostic approach introduced holds promise for future deliberations. The study underscores the significance of integrating LLMs and knowledge processing techniques in scientific research, providing a foundation for advancements in data assimilation and utilization.
38 pages and 10 figures; Fixed figure resolution
References in corpus (10)
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- LoRA: Low-Rank Adaptation of Large Language Models
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- Scalable Extraction of Training Data from (Production) Language Models
- C-Pack: Packed Resources For General Chinese Embeddings
- The False Promise of Imitating Proprietary LLMs
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
- Knowledge Unlearning for LLMs: Tasks, Methods, and Challenges
- Rethinking Machine Unlearning for Large Language Models