3 papers
cs.AI2026
Measuring Activation Control in Large Language Models
Marek Mateusz Kowalski, Joshua Fonseca Rivera, Uzay Macar +1
Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware model…
cs.AI2026
Item Response Theory for AI Safety
Joshua Fonseca Rivera, Neil Shah, David Demitri Africa +1
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because b…
cs.CL2026
Steering Awareness: Detecting Activation Steering from Within
Joshua Fonseca Rivera, David Demitri Africa
Activation steering -- adding a vector to a model's residual stream to modify its behavior -- is widely used in safety evaluations as if the model cannot detect the intervention. W…