6 papers
"Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders
Adelaide Danilov, Aria Nourbakhsh, Oleksandr Marchenko Breneur +1
How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains underexplored. We study speaker repr…
CSULoRA: Closest Safe Update Low-Rank Adaptation
Oleksandr Marchenko Breneur, Adelaide Danilov, Aria Nourbakhsh +1
Low-rank adaptation has become a standard method for parameter-efficient fine-tuning of large language models, but even small amounts of unsafe or adversarial fine-tuning data can…
AEyeDE: An Attention-Based Attribution Framework for AI-Generated Text Detection
Aria Nourbakhsh, Adelaide Danilov, Christoph Schommer +1
Detecting AI-generated text is becoming increasingly challenging as modern language models approach human-level fluency and can evade detectors that rely on surface statistics or l…
Evaluating Explainable AI Attribution Methods in Neural Machine Translation via Attention-Guided Knowledge Distillation
Aria Nourbakhsh, Salima Lamsiyah, Adelaide Danilov +1
The study of the attribution of input features to the output of neural network models is an active area of research. While numerous Explainable AI (XAI) techniques have been propos…
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
Oleksandr Marchenko Breneur, Adelaide Danilov, Aria Nourbakhsh +1
We present NOTAI.AI, an explainable framework for machine-generated text detection that extends Fast-DetectGPT by integrating curvature-based signals with neural and stylometric fe…
Cluster Purge Loss: Structuring Transformer Embeddings for Equivalent Mutants Detection
Adelaide Danilov, Aria Nourbakhsh, Christoph Schommer
Recent pre-trained transformer models achieve superior performance in various code processing objectives. However, although effective at optimizing decision boundaries, common appr…