Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field
arXiv:2412.08258 · doi:10.1016/j.ipm.2025.104262
Abstract
Ontologies of research topics are crucial for structuring scientific knowledge, enabling scientists to navigate vast amounts of research, and forming the backbone of intelligent systems such as search engines and recommendation systems. However, manual creation of these ontologies is expensive, slow, and often results in outdated and overly general representations. As a solution, researchers have been investigating ways to automate or semi-automate the process of generating these ontologies. This paper offers a comprehensive analysis of the ability of large language models (LLMs) to identify semantic relationships between different research topics, which is a critical step in the development of such ontologies. To this end, we developed a gold standard based on the IEEE Thesaurus to evaluate the task of identifying four types of relationships between pairs of topics: broader, narrower, same-as, and other. Our study evaluates the performance of seventeen LLMs, which differ in scale, accessibility (open vs. proprietary), and model type (full vs. quantised), while also assessing four zero-shot reasoning strategies. Several models have achieved outstanding results, including Mixtral-8x7B, Dolphin-Mistral-7B, and Claude 3 Sonnet, with F1-scores of 0.847, 0.920, and 0.967, respectively. Furthermore, our findings demonstrate that smaller, quantised models, when optimised through prompt engineering, can deliver performance comparable to much larger proprietary models, while requiring significantly fewer computational resources.
Now accepted to Information Processing & Management. this is the camera ready
References in corpus (28)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- A Survey of Large Language Models
- BERTopic: Neural topic modeling with a class-based TF-IDF procedure
- Large Language Models are Zero-Shot Reasoners
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models
- LIMA: Less Is More for Alignment
- The Flan Collection: Designing Data and Methods for Effective Instruction Tuning
- Instruction Tuning for Large Language Models: A Survey
- Orca: Progressive Learning from Complex Explanation Traces of GPT-4
- From human experts to machines: An LLM supported approach to ontology and knowledge graph construction
- Orca 2: Teaching Small Language Models How to Reason
- OpenChat: Advancing Open-source Language Models with Mixed-Quality Data
- TIES-Merging: Resolving Interference When Merging Models
- Self-Verification Improves Few-Shot Clinical Information Extraction
- A Survey on Knowledge Organization Systems of Research Fields: Resources and Challenges
- SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning
- Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning
- LLMs4OL: Large Language Models for Ontology Learning
- The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction
- Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
- Advancing LLM Reasoning Generalists with Preference Trees
- Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language Models
- SemOpenAlex: The Scientific Landscape in 26 Billion RDF Triples