9 papers
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
Yidong Wang, Yunze Song, Tingyuan Zhu +11
The adoption of Large Language Models (LLMs) as automated evaluators (LLM-as-a-judge) has revealed critical inconsistencies in current evaluation frameworks. We identify two fundam…
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
Yidong Wang, Xin Wang, Cunxiang Wang +9
Self-Rewarding Language Models propose an architecture in which the Large Language Models(LLMs) both generates responses and evaluates its own outputs via LLM-as-a-Judge prompting,…
RewardAnything: Generalizable Principle-Following Reward Models
Zhuohao Yu, Jiali Zeng, Weizheng Gu +7
Reward Models, essential for guiding Large Language Model optimization, are typically trained on fixed preference datasets, resulting in rigid alignment to single, implicit prefere…
Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark
Jian Wu, Linyi Yang, Zhen Wang +2
While Large Language Models (LLMs) excel in question-answering (QA) tasks, their real reasoning abilities on multiple evidence retrieval and integration on Multi-hop QA tasks remai…
RAGLAB: A Modular and Research-Oriented Unified Framework for Retrieval-Augmented Generation
Xuanwang Zhang, Yunze Song, Yidong Wang +10
Large Language Models (LLMs) demonstrate human-level capabilities in dialogue, reasoning, and knowledge retention. However, even the most advanced LLMs face challenges such as hall…
AutoSurvey: Large Language Models Can Automatically Write Surveys
Yidong Wang, Qi Guo, Wenjin Yao +10
This paper introduces AutoSurvey, a speedy and well-organized methodology for automating the creation of comprehensive literature surveys in rapidly evolving fields like artificial…