4 papers
Test-Time Compute Games
Ander Artola Velasco, Dimitrios Rontogiannis, Stratis Tsirtsis +1
Test-time compute has emerged as a promising strategy to enhance the reasoning abilities of large language models (LLMs). However, this strategy has in turn increased how much user…
Interpretability-by-Design with Accurate Locally Additive Models and Conditional Feature Effects
Vasilis Gkolemis, Loukas Kavouras, Dimitrios Kyriakopoulos +5
Generalized additive models (GAMs) offer interpretability through independent univariate feature effects but underfit when interactions are present in data. GAMs add selected p…
GLANCE: Global Actions in a Nutshell for Counterfactual Explainability
Loukas Kavouras, Eleni Psaroudaki, Konstantinos Tsopelas +9
The widespread deployment of machine learning systems in critical real-world decision-making applications has highlighted the urgent need for counterfactual explainability methods…
Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks
Dimitrios Rontogiannis, Maxime Peyrard, Nicolas Baldwin +3
Standard single-turn, static benchmarks fall short in evaluating the nuanced capabilities of Large Language Models (LLMs) on complex tasks such as software engineering. In this wor…