activity
20242026
collaborators

7 papers

cs.CL2026

Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models

Varvara Arzt, Allan Hanbury, Terra Blevins

We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial language…

cs.IR2026

The LLM Effect on IR Benchmarks: A Meta-Analysis of Effectiveness, Baselines, and Contamination

Moritz Staudinger, Wojciech Kusa, Allan Hanbury

Benchmark collections have long enabled controlled comparison and cumulative progress in Information Retrieval (IR). However, prior meta-analyses have shown that reported effective…

cs.CL2025

Relation Extraction or Pattern Matching? Unravelling the Generalisation Limits of Language Models for Biographical RE

Varvara Arzt, Allan Hanbury, Michael Wiegand +2

Analysing the generalisation capabilities of relation extraction (RE) models is crucial for assessing whether they learn robust relational patterns or rely on spurious correlations…

cs.DL2025

Compare: A Framework for Scientific Comparisons

Moritz Staudinger, Wojciech Kusa, Matteo Cancellieri +3

Navigating the vast and rapidly increasing sea of academic publications to identify institutional synergies, benchmark research contributions and pinpoint key research contribution…

cs.IR2024

A Reproducibility and Generalizability Study of Large Language Models for Query Generation

Moritz Staudinger, Wojciech Kusa, Florina Piroi +2

Systematic literature reviews (SLRs) are a cornerstone of academic research, yet they are often labour-intensive and time-consuming due to the detailed literature curation process.…

cs.CL2024

Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards

Varvara Arzt, Allan Hanbury

This paper investigates the transparency in the creation of benchmarks and the use of leaderboards for measuring progress in NLP, with a focus on the relation extraction (RE) task.…