activity
20192026
collaborators

17 papers

cs.AI2026

MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

Ruoyu Wu, Shenfu Xie, Yinqian Sun +2

Interactive clinical agents operate under partial observability, so reliable care depends on reaching the correct diagnosis through evidence-grounded, safe interactions. Yet existi…

cs.RO2026

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

Mingyang Lyu, Yinqian Sun, Yiyang Jia +5

In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models continue to advance toward gener…

cs.AI2026

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

Linghao Feng, Yinqian Sun, Dongqi Liang +8

Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature analysis to laboratory planning a…

cs.AI2026

ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI

Haibo Tong, Feifei Zhao, Linghao Feng +18

Rapidly evolving AI exhibits increasingly strong autonomy and goal-directed capabilities, accompanied by derivative systemic risks that are more unpredictable, difficult to control…

cs.AI2026

CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models

Haibo Tong, Zeyang Yue, Feifei Zhao +6

Whether Large Language Models (LLMs) truly possess human-like Theory of Mind (ToM) capabilities has garnered increasing attention. However, existing benchmarks remain largely restr…

cs.LG2025

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models

Mingyang Lyu, Yinqian Sun, Erliang Lin +4

Vision-Language-Action (VLA) models such as OpenVLA, Octo, and have shown strong generalization by leveraging large-scale demonstrations, yet their performance is still funda…