14 papers
ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation
Ximo Zhu, Ruiqi Liu, Rong Wang +8
On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local conf…
HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
Zhouyuan Ma, Yutao Wu, Hanxun Huang +6
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is kn…
Internal Safety Collapse in Frontier Large Language Models
Yutao Wu, Xiao Liu, Yifeng Gao +7
This work identifies a critical failure mode in frontier large language models (LLMs), which we term Internal Safety Collapse (ISC): under certain task conditions, models enter a s…
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
Jiayu Li, Yunhan Zhao, Xiang Zheng +4
Vision-Language-Action (VLA) models enable robots to interpret natural-language instructions and perform diverse tasks, yet their integration of perception, language, and control i…
HaluMem: Evaluating Hallucinations in Memory Systems of Agents
Ding Chen, Simin Niu, Kehang Li +6
Memory systems are key components that enable AI systems such as LLMs and AI agents to achieve long-term learning and sustained interaction. However, during memory storage and retr…
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
Zonghuan Xu, Jiayu Li, Yunhan Zhao +3
Vision-Language-Action (VLA) models map multimodal perception and language instructions to executable robot actions, making them particularly vulnerable to behavioral backdoor mani…