Evaluating Large Language Models in Process Mining: Capabilities, Benchmarks, and Evaluation Strategies
arXiv:2403.06749 · doi:10.1007/978-3-031-61007-3_2
Abstract
Using Large Language Models (LLMs) for Process Mining (PM) tasks is becoming increasingly essential, and initial approaches yield promising results. However, little attention has been given to developing strategies for evaluating and benchmarking the utility of incorporating LLMs into PM tasks. This paper reviews the current implementations of LLMs in PM and reflects on three different questions. 1) What is the minimal set of capabilities required for PM on LLMs? 2) Which benchmark strategies help choose optimal LLMs for PM? 3) How do we evaluate the output of LLMs on specific PM tasks? The answer to these questions is fundamental to the development of comprehensive process mining benchmarks on LLMs covering different tasks and implementation paradigms.
References in corpus (17)
- A Survey on Evaluation of Large Language Models
- Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
- Evaluating the Text-to-SQL Capabilities of Large Language Models
- MMBench: Is Your Multi-modal Model an All-around Player?
- YaRN: Efficient Context Window Extension of Large Language Models
- ARB: Advanced Reasoning Benchmark for Large Language Models
- LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
- Review of Large Vision Models and Visual Prompt Engineering
- Leveraging Large Language Models (LLMs) for Process Mining (Technical Report)
- Chit-Chat or Deep Talk: Prompt Engineering for Process Mining
- Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models
- Large Language Models for Automated Open-domain Scientific Hypotheses Discovery
- BAMBOO: A Comprehensive Benchmark for Evaluating Long Text Modeling Capacities of Large Language Models
- The Confidence-Competence Gap in Large Language Models: A Cognitive Study
- Abstractions, Scenarios, and Prompt Definitions for Process Mining with LLMs: A Case Study
- Self-Evaluation Improves Selective Generation in Large Language Models