collaborators

6 papers

cs.CL2026

Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes

Jeremias Ferrao, Niclas Müller-Hof, Iustin Sîrbu +2

We argue that safety classifiers should model user intent as an explicit signal between the prompt and the final label. To study this, we introduce AIMS, a human-annotated dataset…

cs.AI2025

The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features

Jeremias Ferrao, Matthijs van der Lende, Ilija Lichkovski +1

Prevailing alignment methods induce opaque parameter changes, obscuring what models truly learn. To address this, we introduce Feature Steering with Reinforcement Learning (FSRL),…

cs.CL2025

What Really Counts? Examining Step and Token Level Attribution in Multilingual CoT Reasoning

Jeremias Ferrao, Ezgi Basar, Khondoker Ittehadul Islam +1

This study investigates the attribution patterns underlying Chain-of-Thought (CoT) reasoning in multilingual LLMs. While prior works demonstrate the role of CoT prompting in improv…

cs.LG2025

Self-Ablating Transformers: More Interpretability, Less Sparsity

Jeremias Ferrao, Luhan Mikaelson, Keenan Pepper +1

A growing intuition in machine learning suggests a link between sparsity and interpretability. We introduce a novel self-ablation mechanism to investigate this connection ante-hoc…

stat.ML2025

Evaluating Uncertainty in Deep Gaussian Processes

Matthijs van der Lende, Jeremias Lino Ferrao, Niclas Müller-Hof

Reliable uncertainty estimates are crucial in modern machine learning. Deep Gaussian Processes (DGPs) and Deep Sigma Point Processes (DSPPs) extend GPs hierarchically, offering pro…

cs.LG2025

World Model Agents with Change-Based Intrinsic Motivation

Jeremias Ferrao, Rafael Cunha

Sparse reward environments pose a significant challenge for reinforcement learning due to the scarcity of feedback. Intrinsic motivation and transfer learning have emerged as promi…