7 citations · 7 across the 2 of their papers we have counts for
3 papers
cs.CL2025
Evaluating Scoring Bias in LLM-as-a-Judge
Qingquan Li, Shaoyu Dou, Kailai Shao +2
The "LLM-as-a-Judge" paradigm, using Large Language Models (LLMs) as automated evaluators, is pivotal to LLM development, offering scalable feedback for complex tasks. However, the…
cs.CL2024
Beyond Benchmarking: A New Paradigm for Evaluation and Assessment of Large Language Models
Jin Liu, Qingquan Li, Wenlong Du
In current benchmarks for evaluating large language models (LLMs), there are issues such as evaluation content restriction, untimely updates, and lack of optimization guidance. In…
cs.CL2023★ 7 cited
MUSER: A Multi-View Similar Case Retrieval Dataset
Qingquan Li, Yiran Hu, Feng Yao +4
Similar case retrieval (SCR) is a representative legal AI application that plays a pivotal role in promoting judicial fairness. However, existing SCR datasets only focus on the fac…