13 papers
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Yifan Ding, Xincheng Wei, Yoshua Y. Li +7
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stron…
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
Yunhao Feng, Ruixiao Lin, Ming Wen +12
LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed sa…
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
Yunhao Feng, Xiaohu Du, Xinhao Deng +13
Computer-use agents extend language models from text generation to sustained interaction with files, terminals, browsers, and external tools. This shift creates safety risks that a…
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
Yunhao Feng, Yifan Ding, Yingshui Tan +6
Skill-based agent systems tackle complex tasks by composing reusable skills, improving modularity and scalability while introducing a largely unexamined security attack surface. We…
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
Yutao Wu, Xiao Liu, Yinghui Li +5
Knowledge poisoning poses a critical threat to Retrieval-Augmented Generation (RAG) systems by injecting adversarial content into knowledge bases, tricking Large Language Models (L…
SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning
Run Zou, Jianhang Ding, Yifan Ding +3
Instruction tuning has optimized the specialized capabilities of large language models (LLMs), but it often requires extensive datasets and prolonged training times. The challenge…