activity
20242026
collaborators

14 papers

cs.LG2026

ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation

Ximo Zhu, Ruiqi Liu, Rong Wang +8

On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local conf…

cs.CL2026

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

Zhouyuan Ma, Yutao Wu, Hanxun Huang +6

Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is kn…

cs.CL2026

Internal Safety Collapse in Frontier Large Language Models

Yutao Wu, Xiao Liu, Yifeng Gao +7

This work identifies a critical failure mode in frontier large language models (LLMs), which we term Internal Safety Collapse (ISC): under certain task conditions, models enter a s…

cs.CR2025

AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models

Jiayu Li, Yunhan Zhao, Xiang Zheng +4

Vision-Language-Action (VLA) models enable robots to interpret natural-language instructions and perform diverse tasks, yet their integration of perception, language, and control i…

cs.CL2025

HaluMem: Evaluating Hallucinations in Memory Systems of Agents

Ding Chen, Simin Niu, Kehang Li +6

Memory systems are key components that enable AI systems such as LLMs and AI agents to achieve long-term learning and sustained interaction. However, during memory storage and retr…

cs.CR2025

DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models

Zonghuan Xu, Jiayu Li, Yunhan Zhao +3

Vision-Language-Action (VLA) models map multimodal perception and language instructions to executable robot actions, making them particularly vulnerable to behavioral backdoor mani…