2 papers
cs.CL2026
D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios
Siyi Hao, Yidi Cao, Linhao Yu +2
With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer…
cs.CL2026
DEP: A Decentralized Large Language Model Evaluation Protocol
Jianxiang Peng, Junhao Li, Hongxiang Wang +15
With the rapid development of Large Language Models (LLMs), a large number of benchmarks have been proposed. However, most benchmarks lack unified evaluation standard and require t…