Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
Shree Harsha Bokkahalli Satish, Gustav Eje Henter, Ãva Székely
Recent work in benchmarking bias and fairness in speech large language models (SpeechLLMs) has relied heavily on multiple-choice question answering (MCQA) formats. The model is tas…
cs.CL2024
CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models
Ruibo Tu, Hedvig Kjellström, Gustav Eje Henter +1
Causal reasoning capabilities are essential for large language models (LLMs) in a wide range of applications, such as education and healthcare. But there is still a lack of benchma…
cs.CL2024
Exploring Internal Numeracy in Language Models: A Case Study on ALBERT
Ulme Wennberg, Gustav Eje Henter
It has been found that Transformer-based language models have the ability to perform basic quantitative reasoning. In this paper, we propose a method for studying how these models…