2 citations · 3 across the 10 of their papers we have counts for
6 papers · 1 filter
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
Yixu Wang, Xin Wang, Yang Yao +5
The rapid integration of Large Language Models (LLMs) into high-stakes domains necessitates reliable safety and compliance evaluation. However, existing static benchmarks are ill-e…
Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
Yang Yao, Yixu Wang, Yuxuan Zhang +9
As an embodiment of intelligence evolution toward interconnected architectures, Deep Research Agents (DRAs) systematically exhibit the capabilities in task decomposition, cross-sou…
A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
Xingjun Ma, Yixu Wang, Hengyuan Xu +18
The rapid evolution of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has driven major gains in reasoning, perception, and generation across language and…
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45 Law
Shanghai AI Lab, :, Yicheng Bao +115
We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framewo…
The Other Mind: How Language Models Exhibit Human Temporal Cognition
Lingyu Li, Yang Yao, Yixu Wang +3
As Large Language Models (LLMs) continue to advance, they exhibit certain cognitive patterns similar to those of humans that are not directly specified in training data. This study…
Reflection-Bench: Evaluating Epistemic Agency in Large Language Models
Lingyu Li, Yixu Wang, Haiquan Zhao +4
With large language models (LLMs) increasingly deployed as cognitive engines for AI agents, the reliability and effectiveness critically hinge on their intrinsic epistemic agency,…