activity
20242026
most citedMEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts

2 citations · 2 across the 17 of their papers we have counts for

collaborators
Showing cs.AIShow all

10 papers · 1 filter

cs.AI2026

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

Zongrui Wang, Xiangyang Zhu, Sicheng Wang +13

Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for…

cs.AI2026

Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions

Kaicheng Shen, Lingyu Li, Wen Wu +3

AI companions powered by large language models increasingly interact with cognition-developing users, including children and adolescents, creating risks that may accumulate over ti…

cs.AI2026

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience

Xinyang Liao, Lingyu Li, Huacan Liu +5

As Large Language Model based agents enter autonomous scientific research, their ability to resist pseudoscience becomes increasingly important. Otherwise, such systems may rapidly…

cs.AI2026

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence

Xinquan Chen, Zhenyun Yin, Shan He +38

As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool use, and real environment intera…

cs.AI2025

MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs

Liang Shan, Kaicheng Shen, Wen Wu +9

Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address implicit, domain-specific risks. T…

cs.AI2025

Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports

Yang Yao, Yixu Wang, Yuxuan Zhang +9

As an embodiment of intelligence evolution toward interconnected architectures, Deep Research Agents (DRAs) systematically exhibit the capabilities in task decomposition, cross-sou…