most citedTrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2025

Persona-Aware Alignment Framework for Personalized Dialogue Generation

Guanrong Li, Xinyu Liu, Zhen Wu +1

Personalized dialogue generation aims to leverage persona profiles and dialogue history to generate persona-relevant and consistent responses. Mainstream models typically rely on t…

cs.AI20251 cited

TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them

Yidong Wang, Yunze Song, Tingyuan Zhu +11

The adoption of Large Language Models (LLMs) as automated evaluators (LLM-as-a-judge) has revealed critical inconsistencies in current evaluation frameworks. We identify two fundam…

cs.CL2025

RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios

Fei Zhao, Chengqiang Lu, Yufan Shen +9

While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset…

cs.CL2025

PersuasiveToM: A Benchmark for Evaluating Machine Theory of Mind in Persuasive Dialogues

Fangxu Yu, Lai Jiang, Shenyi Huang +2

The ability to understand and predict the mental states of oneself and others, known as the Theory of Mind (ToM), is crucial for effective social scenarios. Although recent studies…

cs.AI2025

Rethinking Relation Extraction: Beyond Shortcuts to Generalization with a Debiased Benchmark

Liang He, Yougang Chu, Zhen Wu +3

Benchmarks are crucial for evaluating machine learning algorithm performance, facilitating comparison and identifying superior solutions. However, biases within datasets can lead m…

cs.CL2024

RAGLAB: A Modular and Research-Oriented Unified Framework for Retrieval-Augmented Generation

Xuanwang Zhang, Yunze Song, Yidong Wang +10

Large Language Models (LLMs) demonstrate human-level capabilities in dialogue, reasoning, and knowledge retention. However, even the most advanced LLMs face challenges such as hall…