Publications (25)
What Matters When Building Universal Multilingual Named Entity Recognition Models?
Jonas Golde, Patrick Haller, Alan Akbik
Recent progress in universal multilingual named entity recognition (NER) has been driven by multilingual transformer models, task-specific architectures, custom loss functions, and…
Familiarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training Data
Jonas Golde, Patrick Haller, Max Ploner +3
Zero-shot named entity recognition (NER) is the task of detecting named entities of specific types (such as 'Person' or 'Medicine') without any training examples. Current research…
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
Patrick Haller, Jonas Golde, Alan Akbik
This paper explores the potential of recurrent neural networks (RNNs) and other subquadratic architectures as competitive alternatives to transformer-based models in low-resource l…
OpinionGPT: Modelling Explicit Biases in Instruction-Tuned LLMs
Patrick Haller, Ansar Aynetdinov, Alan Akbik
Instruction-tuned Large Language Models (LLMs) have recently showcased remarkable ability to generate fitting responses to natural language instructions. However, an open research…
FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
Jonas Golde, Patrick Haller, Alan Akbik
Recent multilingual named entity recognition (NER) work has shown that large language models (LLMs) can provide effective synthetic supervision, yet such datasets have mostly appea…
MastermindEval: A Simple But Scalable Reasoning Benchmark
Jonas Golde, Patrick Haller, Fabio Barth +1
Recent advancements in large language models (LLMs) have led to remarkable performance across a wide range of language understanding and mathematical tasks. As a result, increasing…
ScanDL: A Diffusion Model for Generating Synthetic Scanpaths on Texts
Lena S. Bolliger, David R. Reich, Patrick Haller +3
Eye movements in reading play a crucial role in psycholinguistic research studying the cognitive mechanisms underlying human language processing. More recently, the tight coupling…
Language models emulate certain cognitive profiles: An investigation of how predictability measures interact with individual differences
Patrick Haller, Lena S. Bolliger, Lena A. Jäger
To date, most investigations on surprisal and entropy effects in reading have been conducted on the group level, disregarding individual differences. In this work, we revisit the p…
Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling
Ansar Aynetdinov, Patrick Haller, Alan Akbik
Recent research has shown that filtering massive English web corpora into high-quality subsets significantly improves training efficiency. However, for high-resource non-English la…
PoTeC: A German Naturalistic Eye-tracking-while-reading Corpus
Deborah N. Jakobi, Thomas Kern, David R. Reich +2
The Potsdam Textbook Corpus (PoTeC) is a naturalistic eye-tracking-while-reading corpus containing data from 75 participants reading 12 scientific texts. PoTeC is the first natural…
EMTeC: A Corpus of Eye Movements on Machine-Generated Texts
Lena Sophia Bolliger, Patrick Haller, Isabelle Caroline Rose Cretton +3
The Eye Movements on Machine-Generated Texts Corpus (EMTeC) is a naturalistic eye-movements-while-reading corpus of 107 native English speakers reading machine-generated texts. The…
Eye-tracking based classification of Mandarin Chinese readers with and without dyslexia using neural sequence models
Patrick Haller, Andreas Säuberli, Sarah Elisabeth Kiener +3
Eye movements are known to reflect cognitive processes in reading, and psychological reading research has shown that eye gaze patterns differ between readers with and without dysle…
BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing
Jason Alan Fries, Leon Weber, Natasha Seelam +40
Training and evaluating language models increasingly requires the construction of meta-datasets --diverse collections of curated data with clear provenance. Natural language prompt…
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
BigScience Workshop, :, Teven Le Scao +391
Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to wi…
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
Patrick Haller, Jonas Golde, Alan Akbik
We study architectural and optimization techniques for sample-efficient language modeling under the constraints of the BabyLM 2025 shared task. Our model, BLaLM, replaces self-atte…
Revisiting the Uniform Information Density Hypothesis
Clara Meister, Tiago Pimentel, Patrick Haller +3
The uniform information density (UID) hypothesis posits a preference among language users for utterances structured such that information is distributed uniformly across a signal.…
Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
Jonas Golde, Patrick Haller, Felix Hamborg +2
Most NLP tasks are modeled as supervised learning and thus require labeled training data to train effective models. However, manually producing such data at sufficient quality and…
Eyettention: An Attention-based Dual-Sequence Model for Predicting Human Scanpaths during Reading
Shuwen Deng, David R. Reich, Paul Prasse +3
Eye movements during reading offer insights into both the reader's cognitive processes and the characteristics of the text that is being read. Hence, the analysis of scanpaths in r…
What Matters in Linearizing Language Models? A Comparative Study of Architecture, Scale, and Task Adaptation
Patrick Haller, Jonas Golde, Alan Akbik
Linearization has emerged as a strategy for developing efficient language models (LMs). Starting from an existing Transformer-based LM, linearization replaces the attention compone…
PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
Patrick Haller, Fabio Barth, Jonas Golde +2
Vision-language models (VLMs) have demonstrated remarkable progress in multimodal reasoning. However, existing benchmarks remain limited in terms of high-quality, human-verified ex…
PECC: Problem Extraction and Coding Challenges
Patrick Haller, Jonas Golde, Alan Akbik
Recent advancements in large language models (LLMs) have showcased their exceptional abilities across various tasks, such as code generation, problem-solving and reasoning. Existin…
Leveraging In-Context Learning for Political Bias Testing of LLMs
Patrick Haller, Jannis Vamvas, Rico Sennrich +1
A growing body of work has been querying LLMs with political questions to evaluate their potential biases. However, this probing method has limited stability, making comparisons be…
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
Daniel Christoph, Max Ploner, Patrick Haller +1
Sample efficiency is a crucial property of language models with practical implications for training efficiency. In real-world text, information follows a long-tailed distribution.…
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
Patrick Haller, Mark Ibrahim, Polina Kirichenko +2
For Large Language Models (LLMs) to be reliable, they must learn robust knowledge that can be generally applied in diverse settings -- often unlike those seen during training. Yet,…
Digital Comprehensibility Assessment of Simplified Texts among Persons with Intellectual Disabilities
Andreas Säuberli, Franz Holzknecht, Patrick Haller +4
Text simplification refers to the process of increasing the comprehensibility of texts. Automatic text simplification models are most commonly evaluated by experts or crowdworkers…