7 papers
When Truth Is Distributed: Misinformation Derails Collective Fact Recovery in LLM-Based Multi-Agent Systems
Chenfei Yan, Zeyang Yue, Feifei Zhao +6
LLM-based multi-agent systems promise effective collaborative reasoning, but communication may amplify local errors into collective risks, and while existing evaluations emphasize…
Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
Mingyang Lyu, Yinqian Sun, Erliang Lin +4
Vision-Language-Action (VLA) models such as OpenVLA, Octo, and have shown strong generalization by leveraging large-scale demonstrations, yet their performance is still fund…
CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model
Zeyang Yue, Chenfei Yan, Feifei Zhao +5
Whether Large Language Models (LLMs) exhibit covert psychological manipulation in complex human-AI interactions has garnered increasing safety concerns. However, existing AI safety…
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
Haibo Tong, Feifei Zhao, Linghao Feng +18
Rapidly evolving AI exhibits increasingly strong autonomy and goal-directed capabilities, accompanied by derivative systemic risks that are more unpredictable, difficult to control…
CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models
Haibo Tong, Zeyang Yue, Feifei Zhao +6
Whether Large Language Models (LLMs) truly possess human-like Theory of Mind (ToM) capabilities has garnered increasing attention. However, existing benchmarks remain largely restr…
Building Altruistic and Moral AI Agent with Brain-inspired Emotional Empathy Mechanisms
Feifei Zhao, Hui Feng, Haibo Tong +5
As AI closely interacts with human society, it is crucial to ensure that its behavior is safe, altruistic, and aligned with human ethical and moral values. However, existing resear…