3 papers
cs.CL2026
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
Christopher Schröder, Christopher Schröder, Lukas Gienapp +3
We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes out a large fraction of att…
cs.CL2026
Targeted Syntactic Evaluation of Language Models on Georgian Case Alignment
Daniel Gallagher, Gerhard Heyer
This paper evaluates the performance of transformer-based language models on split-ergative case alignment in Georgian, a particularly rare system for assigning grammatical cases t…
cs.CL2024
Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language Models
Christopher Schröder, Gerhard Heyer
Active learning is an iterative labeling process that is used to obtain a small labeled subset, despite the absence of labeled data, thereby enabling to train a model for supervise…