1 citations · 1 across the 6 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes
Jeremias Ferrao, Niclas Müller-Hof, Iustin Sîrbu +2
We argue that safety classifiers should model user intent as an explicit signal between the prompt and the final label. To study this, we introduce AIMS, a human-annotated dataset…
cs.CL2025
What Really Counts? Examining Step and Token Level Attribution in Multilingual CoT Reasoning
Jeremias Ferrao, Ezgi Basar, Khondoker Ittehadul Islam +1
This study investigates the attribution patterns underlying Chain-of-Thought (CoT) reasoning in multilingual LLMs. While prior works demonstrate the role of CoT prompting in improv…