4 papers · 1 filter
"Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders
Adelaide Danilov, Aria Nourbakhsh, Oleksandr Marchenko Breneur +1
How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains underexplored. We study speaker repr…
AEyeDE: An Attention-Based Attribution Framework for AI-Generated Text Detection
Aria Nourbakhsh, Adelaide Danilov, Christoph Schommer +1
Detecting AI-generated text is becoming increasingly challenging as modern language models approach human-level fluency and can evade detectors that rely on surface statistics or l…
Evaluating Explainable AI Attribution Methods in Neural Machine Translation via Attention-Guided Knowledge Distillation
Aria Nourbakhsh, Salima Lamsiyah, Adelaide Danilov +1
The study of the attribution of input features to the output of neural network models is an active area of research. While numerous Explainable AI (XAI) techniques have been propos…
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
Oleksandr Marchenko Breneur, Adelaide Danilov, Aria Nourbakhsh +1
We present NOTAI.AI, an explainable framework for machine-generated text detection that extends Fast-DetectGPT by integrating curvature-based signals with neural and stylometric fe…