activity
20242026
collaborators
Showing cs.AIShow all

7 papers · 1 filter

cs.AI2026

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +306

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…

cs.AI2026

From Model Scaling to System Scaling: Scaling the Harness in Agentic AI

Shangding Gu

This paper studies the next major bottleneck in agentic AI as system scaling, not only model scaling: the design of auditable, persistent, modular, and verifiable architectures aro…

cs.AI2026

Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents

Ahmad Al-Tawaha, Shangding Gu, Peizhi Niu +2

Safety evaluations of memory-equipped LLM agents typically measure within-task safety: whether an agent completes a single scenario safely, often under adversarial conditions such…

cs.AI20261 cited

AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions

Xianyang Liu, Shangding Gu, Dawn Song

Large language model (LLM)-based agents are increasingly expected to negotiate, coordinate, and transact autonomously, yet existing benchmarks lack principled settings for evaluati…

cs.AI2026

Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity

Yingxuan Yang, Chengrui Qu, Muning Wen +5

LLM-based multi-agent systems (MAS) have emerged as a promising approach to tackle complex tasks that are difficult for individual LLMs. A natural strategy is to scale performance…

cs.AI2025

Agentic Web: Weaving the Next Web with AI Agents

Yingxuan Yang, Mulei Ma, Yuxuan Huang +15

The emergence of AI agents powered by large language models (LLMs) marks a pivotal shift toward the Agentic Web, a new phase of the internet defined by autonomous, goal-driven inte…