activity
20232026
most citedLLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts

1 citations · 1 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CL2026

MemEvoBench: Benchmarking Safety Risks from Memory Misevolution in LLM Agents

Weiwei Xie, Shaoxiong Guo, Fan Zhang +5

Equipping Large Language Models (LLMs) with persistent memory enhances interaction continuity and personalization but introduces new safety risks. Specifically, contaminated or bia…

cs.MA2025

When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms

Qibing Ren, Zhijie Zheng, Jiaxuan Guo +3

In this work, we study the risks of collective financial fraud in large-scale multi-agent systems powered by large language model (LLM) agents. We investigate whether agents can co…

cs.AI2025

When Autonomy Goes Rogue: Preparing for Risks of Multi-Agent Collusion in Social Systems

Qibing Ren, Sitao Xie, Longxuan Wei +4

Recent large-scale events like election fraud and financial scams have shown how harmful coordinated efforts by human groups can be. With the rise of autonomous AI systems, there i…

cs.CV2025

One RL to See Them All: Visual Triple Unified Reinforcement Learning

Yan Ma, Linge Du, Xuyang Shen +7

Reinforcement learning (RL) is becoming an important direction for post-training vision-language models (VLMs), but public training methodologies for unified multimodal RL remain m…

cs.CL2024

LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts

Qibing Ren, Hao Li, Dongrui Liu +7

Safety concerns in large language models (LLMs) have gained significant attention due to their exposure to potentially harmful data during pre-training. In this paper, we identify…

cs.CL2024

CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion

Qibing Ren, Chang Gao, Jing Shao +4

The rapid advancement of Large Language Models (LLMs) has brought about remarkable generative capabilities but also raised concerns about their potential misuse. While strategies l…