most citedUnited Minds or Isolated Agents? Exploring Coordination of LLMs under Cognitive Load Theory

3 citations · 3 across the 2 of their papers we have counts for

collaborators

11 papers

cs.AI20263 cited

United Minds or Isolated Agents? Exploring Coordination of LLMs under Cognitive Load Theory

HaoYang Shang, Xuan Liu, Zi Liang +3

Large Language Models (LLMs) exhibit a notable performance ceiling on complex, multi-faceted tasks. As practitioners increasingly rely on heavy context engineering -- curating intr…

cs.AI2026

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +306

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…

cs.CE2026

Diagon: A Programmable Testbed for AI-Agent Cognitive Labor Markets

Xuan Liu, Haoyang Shang, Haojian Jin

AI agents are emerging as market participants that trade delegated cognitive work with one another on behalf of their users. Each agent can act as both a task poster and a contract…

cs.LG2026

Decision Potential Surface: A Theoretical and Practical Approximation of Large Language Model Decision Boundary

Zi Liang, Zhiyao Wu, Haoyang Shang +5

Decision boundary, the subspace of inputs where a machine learning model assigns equal classification probabilities to two classes, is pivotal in revealing core model properties an…

cs.CY2026

Validated Hypotheses as a Lens for Human-Likeness Evaluation in AI Agents

Xuan Liu, HaoYang Shang, Zizhang Liu +5

We propose using validated behavioral hypotheses as a lens for evaluating human-likeness in LLM-based agents. Our key idea is simple: If an agent is human-like, a population of suc…

cs.LG2026

GSS: Gated Subspace Steering for Selective Memorization Mitigation in LLMs

Xuanqi Zhang, Haoyang Shang, Xiaoxiao Li

Large language models (LLMs) can memorize and reproduce training sequences verbatim -- a tendency that undermines both generalization and privacy. Existing mitigation methods apply…