jailbreak attacks 2adversarial evaluation 1adversarial prompting 1best-of-N search 1code encoding 1model safety 1recovery decoding 1safety guards 1self-check defense 1vision-language models 1
From the 2 of 5 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
Qizhen Lan, Xi Xiao, Xiangchen Guan +4
On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own trajectories. Existing methods allocate th…
cs.AI2026
MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents
Zhisheng Chen, Bingfan Zeng, Bangde Cao +8
Long-horizon agents rely on memory to reuse experiences, yet existing memory systems often assume that evidence can be directly consumed through a fixed representation. This leads…