From the 1 of 11 linked papers with an AI index.
11 papers
Multi-Agent LLMs Fail to Explore Each Other
Hyeong Kyu Choi, Jiatong Li, Wendi Li +2
The paper shows that large language model agents struggle to explore each other in multi-agent settings, leading to poor coordination, and introduces the MACE framework that uses s…
Multi-Head Recurrent Memory Agents
Jiatong Li, Samuel Yeh, Sharon Li
Recurrent memory agents extend LLMs to arbitrarily long contexts by iteratively consolidating input into a fixed-size memory window. Despite their scalability, these agents exhibit…
Generative AI impacts on intra-urban inequality and skill premium in Beijing
Xiliu He, Haoxiang Zhao, Mingyi Ma +6
Generative artificial intelligence (GenAI) is the first automation wave to reach high-cognitive tasks at scale, yet its effects on intra-urban inequality remain largely unknown. Us…
Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models
Zhaoyi Li, Jiatong Li, Gangwei Jiang +3
Chain-of-thought (CoT) reasoning has become the standard paradigm for enabling Large Language Models (LLMs) to solve complex problems. However, recent studies reveal a sharp perfor…
Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities
Changdae Oh, Seongheon Park, To Eun Kim +8
Uncertainty quantification (UQ) for large language models (LLMs) is a key building block for safety guardrails of daily LLM applications. Yet, even as LLM agents are increasingly d…
K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge
Zhikai Li, Jiatong Li, Xuewen Liu +7
The rapid development of visual generative models raises the need for more scalable and human-aligned evaluation methods. While the crowdsourced Arena platforms offer human prefere…