Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Measuring Activation Control in Large Language Models
Marek Mateusz Kowalski, Joshua Fonseca Rivera, Uzay Macar +1
Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware model…
cs.AI2026
Item Response Theory for AI Safety
Joshua Fonseca Rivera, Neil Shah, David Demitri Africa +1
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because b…