The Effect of Sampling Temperature on Problem Solving in Large Language Models
arXiv:2402.05201 · doi:10.18653/v1/2024.findings-emnlp.432
Abstract
In this research study, we empirically investigate the effect of sampling temperature on the performance of Large Language Models (LLMs) on various problem-solving tasks. We created a multiple-choice question-and-answer (MCQA) exam by randomly sampling problems from standard LLM benchmarks. Then, we used nine popular LLMs with five prompt-engineering techniques to solve the MCQA problems while increasing the sampling temperature from 0.0 to 1.6. Despite anecdotal reports to the contrary, our empirical results indicate that changes in temperature from 0.0 to 1.0 do not have a statistically significant impact on LLM performance for problem-solving tasks. In addition, these results appear to generalize across LLMs, prompt-engineering techniques, and problem domains. All code, data, and supplemental materials are available on GitHub at: https://github.com/matthewrenze/jhu-llm-temperature
Cited by in corpus (7)
- A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?
- Supporting Energy Policy Research with Large Language Models
- Large Linguistic Models: Investigating LLMs' metalinguistic abilities
- Evaluating The Performance of Using Large Language Models to Automate Summarization of CT Simulation Orders in Radiation Oncology
- Impact of Label Noise from Large Language Models Generated Annotations on Evaluation of Diagnostic Model Performance
- Adaptive Domain Modeling with Language Models: A Multi-Agent Approach to Task Planning
- WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis