collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2025

FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition

Jonas Golde, Patrick Haller, Alan Akbik

Recent multilingual named entity recognition (NER) work has shown that large language models (LLMs) can provide effective synthetic supervision, yet such datasets have mostly appea…

cs.CL2025

Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements

Patrick Haller, Jonas Golde, Alan Akbik

We study architectural and optimization techniques for sample-efficient language modeling under the constraints of the BabyLM 2025 shared task. Our model, BLaLM, replaces self-atte…

cs.CL2025

Question Decomposition for Retrieval-Augmented Generation

Paul J. L. Ammann, Jonas Golde, Alan Akbik

Grounding large language models (LLMs) in verifiable external sources is a well-established strategy for generating reliable answers. Retrieval-augmented generation (RAG) is one su…

cs.CL2025

What Matters in Linearizing Language Models? A Comparative Study of Architecture, Scale, and Task Adaptation

Patrick Haller, Jonas Golde, Alan Akbik

Linearization has emerged as a strategy for developing efficient language models (LMs). Starting from an existing Transformer-based LM, linearization replaces the attention compone…

cs.CL2025

MastermindEval: A Simple But Scalable Reasoning Benchmark

Jonas Golde, Patrick Haller, Fabio Barth +1

Recent advancements in large language models (LLMs) have led to remarkable performance across a wide range of language understanding and mathematical tasks. As a result, increasing…

cs.CL2024

BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models

Patrick Haller, Jonas Golde, Alan Akbik

This paper explores the potential of recurrent neural networks (RNNs) and other subquadratic architectures as competitive alternatives to transformer-based models in low-resource l…