3 papers
cs.CL2026
LLMs Contain Multitudes: How Deployment Context Reshapes Model-Level Preferences and Values
Filip Trhlik, Aoife O'Flynn, Angela Yu +2
Large language models (LLMs) are increasingly characterised in recent evaluation work as having stable, model-level preference and value systems. However, accompanying robustness c…
cs.CL2026
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
Filip Trhlik, Andrew Caines, Paula Buttery
Pre-trained language models (LMs) have, over the last few years, grown substantially in both societal adoption and training costs. This rapid growth in size has constrained progres…
cs.AI2025
Rethinking AI Cultural Alignment
Michal Bravansky, Filip Trhlik, Fazl Barez
As general-purpose artificial intelligence (AI) systems become increasingly integrated with diverse human communities, cultural alignment has emerged as a crucial element in their…