3 citations · 3 across the 2 of their papers we have counts for
11 papers
United Minds or Isolated Agents? Exploring Coordination of LLMs under Cognitive Load Theory
HaoYang Shang, Xuan Liu, Zi Liang +3
Large Language Models (LLMs) exhibit a notable performance ceiling on complex, multi-faceted tasks. As practitioners increasingly rely on heavy context engineering -- curating intr…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Diagon: A Programmable Testbed for AI-Agent Cognitive Labor Markets
Xuan Liu, Haoyang Shang, Haojian Jin
AI agents are emerging as market participants that trade delegated cognitive work with one another on behalf of their users. Each agent can act as both a task poster and a contract…
Decision Potential Surface: A Theoretical and Practical Approximation of Large Language Model Decision Boundary
Zi Liang, Zhiyao Wu, Haoyang Shang +5
Decision boundary, the subspace of inputs where a machine learning model assigns equal classification probabilities to two classes, is pivotal in revealing core model properties an…
Validated Hypotheses as a Lens for Human-Likeness Evaluation in AI Agents
Xuan Liu, HaoYang Shang, Zizhang Liu +5
We propose using validated behavioral hypotheses as a lens for evaluating human-likeness in LLM-based agents. Our key idea is simple: If an agent is human-like, a population of suc…
GSS: Gated Subspace Steering for Selective Memorization Mitigation in LLMs
Xuanqi Zhang, Haoyang Shang, Xiaoxiao Li
Large language models (LLMs) can memorize and reproduce training sequences verbatim -- a tendency that undermines both generalization and privacy. Existing mitigation methods apply…