most citedUnderstanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity

2 citations · 4 across the 5 of their papers we have counts for

collaborators

8 papers

cs.AI20261 cited

AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions

Xianyang Liu, Shangding Gu, Dawn Song

Large language model (LLM)-based agents are increasingly expected to negotiate, coordinate, and transact autonomously, yet existing benchmarks lack principled settings for evaluati…

cs.AI20262 cited

Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity

Yingxuan Yang, Chengrui Qu, Muning Wen +5

LLM-based multi-agent systems (MAS) have emerged as a promising approach to tackle complex tasks that are difficult for individual LLMs. A natural strategy is to scale performance…

cs.LG2025

AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond

Shangding Gu, Xiaohan Wang, Donghao Ying +9

Rapid advances in multimodal models demand benchmarks that rigorously evaluate understanding and reasoning in safety-critical, dynamic real-world settings. We present AccidentBench…

cs.AI20251 cited

Agentic Web: Weaving the Next Web with AI Agents

Yingxuan Yang, Mulei Ma, Yuxuan Huang +15

The emergence of AI agents powered by large language models (LLMs) marks a pivotal shift toward the Agentic Web, a new phase of the internet defined by autonomous, goal-driven inte…

cs.LG2025

Few-Shot Test-Time Optimization Without Retraining for Semiconductor Recipe Generation and Beyond

Shangding Gu, Donghao Ying, Ming Jin +4

We introduce Model Feedback Learning (MFL), a novel test-time optimization framework for optimizing inputs to pre-trained AI models or deployed hardware systems without requiring a…

cs.RO2025

Safe Continual Domain Adaptation after Sim2Real Transfer of Reinforcement Learning Policies in Robotics

Josip Josifovski, Shangding Gu, Mohammadhossein Malmir +5

Domain randomization has emerged as a fundamental technique in reinforcement learning (RL) to facilitate the transfer of policies from simulation to real-world robotic applications…