3 papers
cs.AI2025
Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
Simon Lermen, Mateusz Dziemian, Natalia Pérez-Campanero Antolín
We demonstrate how AI agents can coordinate to deceive oversight systems using automated interpretability of neural networks. Using sparse autoencoders (SAEs) as our experimental f…
cs.AI2025
Identifying Cooperative Personalities in Multi-agent Contexts through Personality Steering with Representation Engineering
Kenneth J. K. Ong, Lye Jia Jun, Hieu Minh "Jord" Nguyen +2
As Large Language Models (LLMs) gain autonomous capabilities, their coordination in multi-agent settings becomes increasingly important. However, they often struggle with cooperati…
cs.CR2024
CryptoFormalEval: Integrating LLMs and Formal Verification for Automated Cryptographic Protocol Vulnerability Detection
Cristian Curaba, Denis D'Ambrosi, Alessandro Minisini +1
Cryptographic protocols play a fundamental role in securing modern digital infrastructure, but they are often deployed without prior formal verification. This could lead to the ado…