activity
20242026
collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL2026

Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection

Yusser Al Ghussin, Daniil Gurgurov, Tanja Baeumel +3

Sparse autoencoders (SAEs) enable feature-level mechanistic interpretability and activation steering in large language models (LLMs), but SAE-based language control remains unrelia…

cs.CL2026

CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark

Daniil Gurgurov, Yusser Al Ghussin, Tanja Baeumel +5

Understanding and controlling the behavior of large language models (LLMs) is an increasingly important topic in multilingual NLP. Beyond prompting or fine-tuning, , i.e.,~manipula…

cs.CL2025

LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings

Sebastian Sztwiertnia, Felix Friedrich, Kristian Kersting +2

Pre-training decoder-only language models relies on vast amounts of high-quality data, yet the availability of such data is increasingly reaching its limits. While metadata is comm…

cs.CL2025

Measuring and Guiding Monosemanticity

Ruben Härle, Felix Friedrich, Manuel Brack +4

There is growing interest in leveraging mechanistic interpretability and controllability to better understand and influence the internal dynamics of large language models (LLMs). H…

cs.CL2025

CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models

Niharika Hegde, Subarnaduti Paul, Lars Joel-Frey +4

Large language models (LLMs) excel at operating at scale by leveraging social media and various data crawled from the web. Whereas existing corpora are diverse, their frequent lack…

cs.CL2025

Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench

Felix Friedrich, Thiemo Ganesha Welsch, Manuel Brack +2

Current diversification strategies for text-to-image (T2I) models often ignore contextual appropriateness, leading to over-diversification where demographic attributes are modified…