Using Large Language Models for Solving Thermodynamic Problems
arXiv:2502.05195 · doi:10.1016/j.compchemeng.2025.109333
Abstract
Large Language Models (LLMs) have made significant progress in reasoning, demonstrating their capability to generate human-like responses. This study analyzes the problem-solving capabilities of LLMs in the domain of thermodynamics. A benchmark of 22 thermodynamic problems to evaluate LLMs is presented that contains both simple and advanced problems. Five different LLMs are assessed: GPT-3.5, GPT-4, and GPT-4o from OpenAI, Llama 3.1 from Meta, and le Chat from MistralAI. The answers of these LLMs were evaluated by trained human experts, following a methodology akin to the grading of academic exam responses. The scores and the consistency of the answers are discussed, together with the analytical skills of the LLMs. Both strengths and weaknesses of the LLMs become evident. They generally yield good results for the simple problems, but also limitations become clear: The LLMs do not provide consistent results, they often fail to fully comprehend the context and make wrong assumptions. Given the complexity and domain-specific nature of the problems, the statistical language modeling approach of the LLMs struggles with the accurate interpretation and the required reasoning. The present results highlight the need for more systematic integration of thermodynamic knowledge with LLMs, for example, by using knowledge-based methods.
This document is the unedited Author's version of a Submitted Work to Computers and Chemical Engineering. 19 pages, 2 figures, SI available (15 pages)
References in corpus (17)
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
- The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels
- Exploring ChatGPT and its Impact on Society
- Performance of ChatGPT on the Test of Understanding Graphs in Kinematics
- A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
- Semantic interoperability and characterization of data provenance in computational molecular engineering
- Physics-based simulation ontology: an ontology to support modelling and reuse of data for physics-based simulation
- System 2 thinking in OpenAI's o1-preview model: Near-perfect performance on a mathematics exam
- Assessing Confidence in AI-Assisted Grading of Physics Exams through Psychometrics: An Exploratory Study
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
- Multi-Perspective Stance Detection
- KnowTD-An Actionable Knowledge Representation System for Thermodynamics
- Towards Ontology-Enhanced Representation Learning for Large Language Models
- What's in an embedding? Would a rose by any embedding smell as sweet?
- Enhancing LLM Problem Solving with REAP: Reflection, Explicit Problem Deconstruction, and Advanced Prompting