1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2026★ 1 cited
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
Thomas Winninger, Boussad Addad, Katarzyna Kapusta
Traditional white-box methods for creating adversarial perturbations against LLMs typically rely only on gradient computation from the targeted model, ignoring the internal mechani…
cs.AI2026
Fast Multi-dimensional Refusal Subspaces via RFM-AGOP
Thomas Winninger
Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behaviours are encoded along sing…
cs.AI2026
Steerability via constraints: a substrate for scalable oversight of coding agents
Thomas Winninger
Coding agents are capable; human oversight is the bottleneck. Unconstrained agents introduce security risks, erode codebase scalability, and make human review increasingly costly.…