1 citations · 1 across the 3 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Fast Multi-dimensional Refusal Subspaces via RFM-AGOP
Thomas Winninger
Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behaviours are encoded along sing…
cs.AI2026
Steerability via constraints: a substrate for scalable oversight of coding agents
Thomas Winninger
Coding agents are capable; human oversight is the bottleneck. Unconstrained agents introduce security risks, erode codebase scalability, and make human review increasingly costly.…