activity
20242026
collaborators

15 papers

cs.AI2026

MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

Ruoyu Wu, Shenfu Xie, Yinqian Sun +2

Interactive clinical agents operate under partial observability, so reliable care depends on reaching the correct diagnosis through evidence-grounded, safe interactions. Yet existi…

cs.CL2026

Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety

Ping Wu, Haibo Tong, Feifei Zhao +7

Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refus…

cs.MA2026

When Truth Is Distributed: Misinformation Derails Collective Fact Recovery in LLM-Based Multi-Agent Systems

Chenfei Yan, Zeyang Yue, Feifei Zhao +6

LLM-based multi-agent systems promise effective collaborative reasoning, but communication may amplify local errors into collective risks, and while existing evaluations emphasize…

cs.AI2026

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

Linghao Feng, Yinqian Sun, Dongqi Liang +8

Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature analysis to laboratory planning a…

cs.AI2026

ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

Lu Jia, Haibo Tong, Feifei Zhao +5

Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments,…

cs.AI2026

CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model

Zeyang Yue, Chenfei Yan, Feifei Zhao +5

Whether Large Language Models (LLMs) exhibit covert psychological manipulation in complex human-AI interactions has garnered increasing safety concerns. However, existing AI safety…