1 citations · 1 across the 16 of their papers we have counts for
4 papers · 1 filter
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
Jiaqiang Li, Yajie Yang, Zhiheng Xi +15
Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-s…
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
Yunhao Feng, Ruixiao Lin, Ming Wen +12
LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed sa…
Mirror: A Multi-Agent System for AI-Assisted Ethics Review
Yifan Ding, Yuhui Shi, Zhiyan Li +10
Ethics review is a foundational mechanism of modern research governance, yet contemporary systems face increasing strain as ethical risks arise as structural consequences of large-…
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
Yixu Wang, Xin Wang, Yang Yao +5
The rapid integration of Large Language Models (LLMs) into high-stakes domains necessitates reliable safety and compliance evaluation. However, existing static benchmarks are ill-e…