6 papers
Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes
Jeremias Ferrao, Niclas Müller-Hof, Iustin Sîrbu +2
We argue that safety classifiers should model user intent as an explicit signal between the prompt and the final label. To study this, we introduce AIMS, a human-annotated dataset…
The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features
Jeremias Ferrao, Matthijs van der Lende, Ilija Lichkovski +1
Prevailing alignment methods induce opaque parameter changes, obscuring what models truly learn. To address this, we introduce Feature Steering with Reinforcement Learning (FSRL),…
What Really Counts? Examining Step and Token Level Attribution in Multilingual CoT Reasoning
Jeremias Ferrao, Ezgi Basar, Khondoker Ittehadul Islam +1
This study investigates the attribution patterns underlying Chain-of-Thought (CoT) reasoning in multilingual LLMs. While prior works demonstrate the role of CoT prompting in improv…
Self-Ablating Transformers: More Interpretability, Less Sparsity
Jeremias Ferrao, Luhan Mikaelson, Keenan Pepper +1
A growing intuition in machine learning suggests a link between sparsity and interpretability. We introduce a novel self-ablation mechanism to investigate this connection ante-hoc…
Evaluating Uncertainty in Deep Gaussian Processes
Matthijs van der Lende, Jeremias Lino Ferrao, Niclas Müller-Hof
Reliable uncertainty estimates are crucial in modern machine learning. Deep Gaussian Processes (DGPs) and Deep Sigma Point Processes (DSPPs) extend GPs hierarchically, offering pro…
World Model Agents with Change-Based Intrinsic Motivation
Jeremias Ferrao, Rafael Cunha
Sparse reward environments pose a significant challenge for reinforcement learning due to the scarcity of feedback. Intrinsic motivation and transfer learning have emerged as promi…