most citedTesting and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness

1 citations · 2 across the 5 of their papers we have counts for

collaborators

7 papers

cs.SE2025

Understanding and Bridging the Planner-Coder Gap: A Systematic Study on the Robustness of Multi-Agent Systems for Code Generation

Zongyi Lyu, Songqiang Chen, Zhenlan Ji +5

Multi-agent systems (MASs) have emerged as a promising paradigm for automated code generation, demonstrating impressive performance on established benchmarks. Despite their prosper…

cs.SE2025

Digging Into the Internal: Causality-Based Analysis of LLM Function Calling

Zhenlan Ji, Daoyuan Wu, Wenxuan Wang +3

Function calling (FC) has emerged as a powerful technique for facilitating large language models (LLMs) to interact with external systems and perform structured tasks. However, the…

cs.CR2025

IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems

Liwen Wang, Wenxuan Wang, Shuai Wang +5

The rapid advancement of Large Language Models (LLMs) has led to the emergence of Multi-Agent Systems (MAS) to perform complex tasks through collaboration. However, the intricate n…

cs.CR2025

SoK: Evaluating Jailbreak Guardrails for Large Language Models

Xunguang Wang, Zhenlan Ji, Wenxuan Wang +3

Large Language Models (LLMs) have achieved remarkable progress, but their deployment has exposed critical vulnerabilities, particularly to jailbreak attacks that circumvent safety…

cs.CV2025

VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models

Bingrui Sima, Linhua Cong, Wenxuan Wang +1

The emergence of Multimodal Large Language Models (MLRMs) has enabled sophisticated visual reasoning capabilities by integrating reinforcement learning and Chain-of-Thought (CoT) s…

cs.CL20251 cited

STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models

Xunguang Wang, Wenxuan Wang, Zhenlan Ji +4

Large Language Models (LLMs) have become increasingly vulnerable to jailbreak attacks that circumvent their safety mechanisms. While existing defense methods either suffer from ada…