2 papers
cs.AI2026
Benchmarking Deception Probes via Black-to-White Performance Boosts
Avi Parrack, Carlo Leonardo Attubato, Stefan Heimersheim
AI assistants will occasionally respond deceptively to user queries. Recently, linear classifiers (called "deception probes") have been trained to distinguish the internal activati…
cs.CL2025
Probing and Steering Evaluation Awareness of Language Models
Jord Nguyen, Khiem Hoang, Carlo Leonardo Attubato +1
Language models can distinguish between testing and deployment phases -- a capability known as evaluation awareness. This has significant safety and policy implications, potentiall…