activity
20242026
collaborators

8 papers

cs.CL2026

ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts

Ronja Schwarz, Jannik Strötgen, Jannik Strötgen

The automatic structural analysis of legal texts is a cornerstone of legal technology, yet the extraction of their logical components remains a significant challenge. In this paper…

cs.CL2026

TalkTag: Fine-Grained Morphosyntactic Error Annotation for Transcribed Speech

Shamira Venturini, Oliver Hennhöfer, Steffen Kinkel +1

Fine-grained morphosyntactic error annotation is important in clinical and developmental language research, yet it is labour-intensive, expert-dependent, and difficult to scale. We…

cs.CL2025

Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes

Mingyang Wang, Lukas Lange, Heike Adel +3

Reasoning language models (RLMs) excel at complex tasks by leveraging a chain-of-thought process to generate structured intermediate steps. However, language mixing, i.e., reasonin…

cs.CL2025

Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models

Mingyang Wang, Heike Adel, Lukas Lange +4

Multilingual language models (MLMs) store factual knowledge across languages but often struggle to provide consistent responses to semantically equivalent prompts in different lang…

cs.CL2025

Bring Your Own Knowledge: A Survey of Methods for LLM Knowledge Expansion

Mingyang Wang, Alisa Stoll, Lukas Lange +3

Adapting large language models (LLMs) to new and diverse knowledge is essential for their lasting effectiveness in real-world applications. This survey provides an overview of stat…

cs.CL2024

Better Call SAUL: Fluent and Consistent Language Model Editing with Generation Regularization

Mingyang Wang, Lukas Lange, Heike Adel +2

To ensure large language models contain up-to-date knowledge, they need to be updated regularly. However, model editing is challenging as it might also affect knowledge that is unr…