8 papers
ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts
Ronja Schwarz, Jannik Strötgen, Jannik Strötgen
The automatic structural analysis of legal texts is a cornerstone of legal technology, yet the extraction of their logical components remains a significant challenge. In this paper…
TalkTag: Fine-Grained Morphosyntactic Error Annotation for Transcribed Speech
Shamira Venturini, Oliver Hennhöfer, Steffen Kinkel +1
Fine-grained morphosyntactic error annotation is important in clinical and developmental language research, yet it is labour-intensive, expert-dependent, and difficult to scale. We…
Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes
Mingyang Wang, Lukas Lange, Heike Adel +3
Reasoning language models (RLMs) excel at complex tasks by leveraging a chain-of-thought process to generate structured intermediate steps. However, language mixing, i.e., reasonin…
Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models
Mingyang Wang, Heike Adel, Lukas Lange +4
Multilingual language models (MLMs) store factual knowledge across languages but often struggle to provide consistent responses to semantically equivalent prompts in different lang…
Bring Your Own Knowledge: A Survey of Methods for LLM Knowledge Expansion
Mingyang Wang, Alisa Stoll, Lukas Lange +3
Adapting large language models (LLMs) to new and diverse knowledge is essential for their lasting effectiveness in real-world applications. This survey provides an overview of stat…
Better Call SAUL: Fluent and Consistent Language Model Editing with Generation Regularization
Mingyang Wang, Lukas Lange, Heike Adel +2
To ensure large language models contain up-to-date knowledge, they need to be updated regularly. However, model editing is challenging as it might also affect knowledge that is unr…