◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Thomas Winninger

3 papers hereh-index 18 citations3 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author2
  • first author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.AI2
  • cs.LG1

identity via Semantic Scholar / OpenAlex

most citedUsing Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models

1 citations · 1 across the 3 of their papers we have counts for

collaborators

3 papers

cs.LG2026★ 1 cited

Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models

Thomas Winninger, Boussad Addad, Katarzyna Kapusta

Traditional white-box methods for creating adversarial perturbations against LLMs typically rely only on gradient computation from the targeted model, ignoring the internal mechani…

cs.AI2026

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

Thomas Winninger

Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behaviours are encoded along sing…

cs.AI2026

Steerability via constraints: a substrate for scalable oversight of coding agents

Thomas Winninger

Coding agents are capable; human oversight is the bottleneck. Unconstrained agents introduce security risks, erode codebase scalability, and make human review increasingly costly.…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.