5 papers
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
GSS: Gated Subspace Steering for Selective Memorization Mitigation in LLMs
Xuanqi Zhang, Haoyang Shang, Xiaoxiao Li
Large language models (LLMs) can memorize and reproduce training sequences verbatim -- a tendency that undermines both generalization and privacy. Existing mitigation methods apply…
Love First, Know Later: Persona-Based Romantic Compatibility Through LLM Text World Engines
Haoyang Shang, Zhengyang Yan, Xuan Liu
We propose Love First, Know Later: a paradigm shift in computational matching that simulates interactions first, then assesses compatibility. Instead of comparing static profiles,…
Mutual Wanting in Human--AI Interaction: Empirical Evidence from Large-Scale Analysis of GPT Model Transitions
HaoYang Shang, Xuan Liu
The rapid evolution of large language models (LLMs) creates complex bidirectional expectations between users and AI systems that are poorly understood. We introduce the concept of…
CoBRA: Programming Cognitive Bias in Social Agents Using Classic Social Science Experiments
Xuan Liu, Haoyang Shang, Haojian Jin
This paper introduces CoBRA, a novel toolkit for systematically specifying agent behavior in LLM-based social simulation. We found that conventional approaches that specify agent b…