4 papers
Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods
Michal Moshkovitz, Suraj Srinivas, Lesia Semenova +7
Despite the proliferation of Explainable AI (XAI) techniques -- from feature attributions to sparse autoencoders -- explanations rarely influence real-world workflows. In practice,…
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
Chung-Yiu Yau, Dawei Li, Athanasios Glentis +3
Lookahead-based acceleration methods, such as Nesterov's momentum, are widely used in optimization, but they often become unreliable in deep learning training mainly due to stochas…
How Much Can We Forget about Data Contamination?
Sebastian Bordt, Suraj Srinivas, Valentyn Boreiko +1
The leakage of benchmark data into the training data has emerged as a significant challenge for evaluating the capabilities of large language models (LLMs). In this work, we challe…
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
Valentyn Boreiko, Alexander Panfilov, Vaclav Voracek +2
A plethora of jailbreaking attacks have been proposed to obtain harmful responses from safety-tuned LLMs. These methods largely succeed in coercing the target output in their origi…