3 papers
cs.CL2026
Don't Ignore the Tail: Decoupling top-K Probabilities for Efficient Language Model Distillation
Sayantan Dasgupta, Trevor Cohn, Timothy Baldwin
The core learning signal used in language model distillation is the standard Kullback-Leibler (KL) divergence between the student and teacher distributions. Traditional KL divergen…
cs.CL2025
Benchmarking Gender and Political Bias in Large Language Models
Jinrui Yang, Xudong Han, Timothy Baldwin
We introduce EuroParlVote, a novel benchmark for evaluating large language models (LLMs) in politically sensitive contexts. It links European Parliament debate speeches to roll-cal…
cs.IR2025
Language Bias in Information Retrieval: The Nature of the Beast and Mitigation Methods
Jinrui Yang, Fan Jiang, Timothy Baldwin
Language fairness in multilingual information retrieval (MLIR) systems is crucial for ensuring equitable access to information across diverse languages. This paper sheds light on t…