24 papers
Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
Ping Wu, Haibo Tong, Feifei Zhao +7
Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refus…
When Truth Is Distributed: Misinformation Derails Collective Fact Recovery in LLM-Based Multi-Agent Systems
Chenfei Yan, Zeyang Yue, Feifei Zhao +6
LLM-based multi-agent systems promise effective collaborative reasoning, but communication may amplify local errors into collective risks, and while existing evaluations emphasize…
ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models
Mingyang Lyu, Yinqian Sun, Yiyang Jia +5
In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models continue to advance toward gener…
Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
Mingyang Lyu, Yinqian Sun, Erliang Lin +4
Vision-Language-Action (VLA) models such as OpenVLA, Octo, and have shown strong generalization by leveraging large-scale demonstrations, yet their performance is still fund…
SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety
Linghao Feng, Yinqian Sun, Dongqi Liang +8
Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature analysis to laboratory planning a…
ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
Lu Jia, Haibo Tong, Feifei Zhao +5
Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments,…