activity
20242026
collaborators

9 papers

cs.CR2026

Stealing Reasoning Traces from Proprietary LLM APIs

Alexander Panfilov, David Schmotz, Ilia Shumailov +5

Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather…

cs.AI2026

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

Joachim Schaeffer, Thomas Jiralerspong, Alexander Panfilov +4

AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untru…

cs.AI2026

Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad +3

An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework for deploying capable but unt…

cs.LG2025

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Luca Baroni, Galvin Khara, Joachim Schaeffer +2

Layer-wise normalization (LN) is an essential component of virtually all transformer-based large language models. While its effects on training stability are well documented, its r…

eess.SY2025

Diagnostic-free onboard battery health assessment

Yunhong Che, Vivek N. Lam, Jinwook Rhyu +5

Diverse usage patterns induce complex and variable aging behaviors in lithium-ion batteries, complicating accurate health diagnosis and prognosis. Separate diagnostic cycles are of…

cs.LG2024

Interpretation of High-Dimensional Regression Coefficients by Comparison with Linearized Compressing Features

Joachim Schaeffer, Jinwook Rhyu, Robin Droop +2

Linear regression is often deemed inherently interpretable; however, challenges arise for high-dimensional data. We focus on further understanding how linear regression approximate…