collaborators

5 papers

cs.LG2026

Beyond Multiple Choice: Evaluating Steering Vectors for Summarization

Joschka Braun, Carsten Eickhoff, Seyed Ali Bahrainian

Steering vectors are a lightweight method for controlling text properties by adding a learned bias to language model activations at inference time. While predominantly studied for…

cs.LG2026

Exploration Hacking: Can LLMs Learn to Resist RL Training?

Eyon Jang, Damon Falck, Joschka Braun +6

Reinforcement learning (RL) has become essential to the post-training of large language models (LLMs) for reasoning, agentic capabilities and alignment. Successful RL relies on suf…

cs.CL2026

Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations

Joschka Braun

Steering vectors are a lightweight method for controlling language model behavior by adding a learned bias to the activations at inference time. Although effective on average, stee…

cs.LG2025

Logit Reweighting for Topic-Focused Summarization

Joschka Braun, Bálint Mucsányi, Seyed Ali Bahrainian

Generating abstractive summaries that adhere to a specific topic remains a significant challenge for language models. While standard approaches, such as fine-tuning, are resource-i…

cs.LG2025

Understanding (Un)Reliability of Steering Vectors in Language Models

Joschka Braun, Carsten Eickhoff, David Krueger +2

Steering vectors are a lightweight method to control language model behavior by adding a learned bias to the activations at inference time. Although steering demonstrates promising…