4 papers
MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems
Mao-xun Huang, Jerry Wang, Yi-Cheng Lai +3
Large language model-based multi-agent systems improve complex problem solving through task decomposition, agent specialization, information exchange, and intermediate validation.…
OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization
Jerry Wang, Hsin-Ling Hsu, Yi-Cheng Lai +2
Production LLMs increasingly rely on toxicity-based moderation filters as a primary defense, assuming that harmful intent correlates with toxic surface wording. We show this assump…
Observer, Not Player: Simulating Theory of Mind in LLMs through Game Observation
Jerry Wang, Ting Yiu Liu
We present an interactive framework for evaluating whether large language models (LLMs) exhibit genuine "understanding" in a simple yet strategic environment. As a running example,…
DeRAG: Black-box Adversarial Attacks on Multiple Retrieval-Augmented Generation Applications via Prompt Injection
Jerry Wang, Fang Yu
Adversarial prompt attacks can significantly alter the reliability of Retrieval-Augmented Generation (RAG) systems by re-ranking them to produce incorrect outputs. In this paper, w…