5 papers
Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
Qizhen Lan, Xi Xiao, Xiangchen Guan +4
On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own trajectories. Existing methods allocate th…
Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks
Haoyu Zhang, Xiangchen Guan, Shibo Zheng +2
We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrela…
Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks
Haoyu Zhang, Zhuoxi Wang, Shibo Zheng +7
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encod…
MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents
Zhisheng Chen, Bingfan Zeng, Bangde Cao +8
Long-horizon agents rely on memory to reuse experiences, yet existing memory systems often assume that evidence can be directly consumed through a fixed representation. This leads…
Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses
Haoyu Zhang, Shibo Zheng, Xiangchen Guan +4
A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% defense success rate. We show it…