5 papers
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
Tiankai Yang, Jiate Li, Yi Nian +5
LLM-based agents increasingly operate across repeated sessions, maintaining task states to ensure continuity. In many deployments, a single agent serves multiple users within a tea…
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
Yucheng Chu, Haoyu Han, Shen Dong +6
Automated short answer grading (ASAG) is critical for scaling educational assessment, yet large language models (LLMs) often struggle with hallucinations and strict rubric adherenc…
PEAR: Planner-Executor Agent Robustness Benchmark
Shen Dong, Mingxuan Zhang, Pengfei He +4
Large Language Model (LLM)-based Multi-Agent Systems (MAS) have emerged as a powerful paradigm for tackling complex, multi-step tasks across diverse domains. However, despite their…
Memory Injection Attacks on LLM Agents via Query-Only Interaction
Shen Dong, Shaochen Xu, Pengfei He +5
Agents powered by large language models (LLMs) have demonstrated strong capabilities in a wide range of complex, real-world applications. However, LLM agents with a compromised mem…
Red-Teaming LLM Multi-Agent Systems via Communication Attacks
Pengfei He, Yupin Lin, Shen Dong +3
Large Language Model-based Multi-Agent Systems (LLM-MAS) have revolutionized complex problem-solving capability by enabling sophisticated agent collaboration through message-based…