4 papers
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Hiskias Dingeto
Natural-language autoencoders score explanations of hidden activations by reconstruction. An explanation is deemed faithful if the activation can be regenerated from it. The test i…
AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations
Hiskias Dingeto, William Leeney
Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed th…
When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs
Hiskias Dingeto, Taeyoun Kwon, Dasol Choi +5
As large language models (LLMs) become increasingly integrated into daily life, audio has emerged as a key interface for human-AI interaction. However, this convenience also introd…
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
Siddhant Panpatil, Hiskias Dingeto, Haon Park
Despite significant advances in alignment techniques, we demonstrate that state-of-the-art language models remain vulnerable to carefully crafted conversational scenarios that can…