4 papers
Recognize Your Orchestrator: An Entropy Dynamics Perspective for LLM Multi-Agent Systems
Junze Zhu, Weihao Chen, Xuanwang Zhang +2
The transition from single-turn models to Multi-Agent Systems (MAS) promises enhanced problem-solving capabilities, yet the centralized orchestration topology remains a critical po…
WebNavigator: Global Web Navigation via Interaction Graph Retrieval
Xuanwang Zhang, Yuteng Han, Jinnan Qi +3
Despite significant advances in autonomous web navigation, current methods remain far from human-level performance in complex web environments. We argue that this limitation stems…
EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations
Xinyun Zhou, Xinfeng Li, Yinan Peng +9
Retrieval-Augmented Generation (RAG) systems are increasingly central to robust AI, enhancing large language model (LLM) faithfulness by incorporating external knowledge. However,…
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
Yidong Wang, Yunze Song, Tingyuan Zhu +11
The adoption of Large Language Models (LLMs) as automated evaluators (LLM-as-a-judge) has revealed critical inconsistencies in current evaluation frameworks. We identify two fundam…