9 papers
MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer
Xuefei Wang, Jialu Wang, Fengbo Zhang +6
Multi-agent systems (MAS) powered by large language models (LLMs) have emerged as a powerful paradigm for complex problem solving, where performance critically depends on the under…
Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior
Luyang Zhang, Jialu Wang, Fei Xue +1
Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on the models producing measurably different…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
When Attribution Patching Lies: Diagnosis and a Second-Order Correction
Luyang Zhang, Jialu Wang
A central goal of mechanistic interpretability is to identify which internal components causally drive a language model's behavior. Because these importance estimates serve as the…
UniPinRec: Unifying Generative Retrieval and Ranking at Pinterest Scale
Hanyu Li, Yi-Ping Hsu, Aditya Mantha +17
Modern recommendation systems predominantly train retrieval and ranking as separate models despite both increasingly relying on large transformers encoding the same user behavior d…
Do Agents Repair When Challenged -- or Just Reply? Challenge, Repair, and Public Correction in a Deployed Agent Forum
Luyang Zhang, Yi-Yun Chu, Jialu Wang +2
As large language model (LLM) agents are deployed in public interactive settings, a key question is whether their communities can sustain challenge, repair, and public correction,…