From the 2 of 5 linked papers with an AI index.
5 papers
Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
Qizhen Lan, Xi Xiao, Xiangchen Guan +4
On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own trajectories. Existing methods allocate th…
MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents
Zhisheng Chen, Bingfan Zeng, Bangde Cao +8
Long-horizon agents rely on memory to reuse experiences, yet existing memory systems often assume that evidence can be directly consumed through a fixed representation. This leads…
Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks
Haoyu Zhang, Xiangchen Guan, Shibo Zheng +2
We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrela…
Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses
Haoyu Zhang, Shibo Zheng, Xiangchen Guan +4
The paper demonstrates that a self‑check defense (SAGE) for language models can be bypassed by combining a code‑completion encoding attack with a best‑of‑N search, dramatically inc…
Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks
Haoyu Zhang, Zhuoxi Wang, Shibo Zheng +7
The paper proposes a guard‑agnostic recovery‑and‑decode module that transcribes encoded or visual text into plain language before applying existing safety classifiers for vision‑la…