9 papers
ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration
Guo Chen, Ziwen Li, Reed Li +4
Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common basis for evaluation across me…
Training Documents Reranker with Search Rubrics for Deep Research Agent
Wenhan Liu, Yu Lu, Qiaolin Xia +8
Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typically select documents through relevance…
OPOD: On-Policy Omni Distillation
Tong Zhao, Yuyang Hu, Reed Li +5
Omni-modal models provide a unified interface for text, images, and audio. However, improving these abilities together remains difficult, as post-training on pooled multimodal data…
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking
Kailin Jiang, Lei Liu, Jian Xi +8
As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evalua…
Reason Only When Needed: Efficient Generative Reward Modeling via Model-Internal Uncertainty
Chao Xue, Yao Wang, Mengqiao Liu +11
Recent advancements in the Generative Reward Model (GRM) have demonstrated its potential to enhance the reasoning abilities of LLMs through Chain-of-Thought (CoT) prompting. Despit…
Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language Models
Chao Xue, Yao Wang, Mengqiao Liu +11
Supervised Fine-Tuning (SFT) is the standard approach for adapting large language models (LLMs) to downstream tasks. However, we observe a persistent failure mode: even after conve…