activity
20242026
collaborators

21 papers

cs.CV2026

Finding DoRI: Discovery of Retained Images in Diffusion Models

Antoni Kowalczuk, Dominik Hintersdorf, Lukas Struppek +3

Text-to-image diffusion models (DMs) have achieved remarkable success in image generation. However, concerns about data privacy and intellectual property remain due to their potent…

cs.CV2026

No Safe Dose: How Training Data Drives Unsafe Image Generation

Felix Friedrich, Lukas Helff, Niharika Hegde +2

Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remains unclear whether and how t…

cs.AI2026

SLR: Automated Synthesis for Scalable Logical Reasoning

Lukas Helff, Ahmad Omar, Felix Friedrich +7

We introduce SLR, an end-to-end framework for systematic evaluation and training of Large Language Models (LLMs) via Scalable Logical Reasoning. Given a user's task specification,…

cs.CL2026

EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection

Christoph Schuhmann, Robert Kaczmarczyk, Gollam Rabby +6

Speech emotion recognition (SER) systems are constrained by existing datasets that typically cover only 6-10 basic emotions, lack scale and diversity, and face ethical challenges w…

cs.CL2025

LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings

Sebastian Sztwiertnia, Felix Friedrich, Kristian Kersting +2

Pre-training decoder-only language models relies on vast amounts of high-quality data, yet the availability of such data is increasingly reaching its limits. While metadata is comm…

cs.CL2025

Measuring and Guiding Monosemanticity

Ruben Härle, Felix Friedrich, Manuel Brack +4

There is growing interest in leveraging mechanistic interpretability and controllability to better understand and influence the internal dynamics of large language models (LLMs). H…