activity
20242026
collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

Junjie Chen, Yuxi Dong, Haitao Li +6

As large language models (LLMs) are increasingly used for long-form generation, reliably evaluating long-form outputs has become a critical challenge. LLM-as-a-judge offers a scala…

cs.CL2026

LegalOne: A Family of Foundation Models for Reliable Legal Reasoning

Haitao Li, Yifan Chen, Shuo Miao +13

While Large Language Models (LLMs) have demonstrated impressive general capabilities, their direct application in the legal domain is often hindered by a lack of precise domain kno…

cs.CL2025

Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation Evaluation

Junjie Chen, Weihang Su, Zhumin Chu +9

The rapid development of large language models (LLMs) has highlighted the need for efficient and reliable methods to evaluate their performance. Traditional evaluation methods ofte…

cs.CL2025

Overview of the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task

Junjie Chen, Haitao Li, Zhumin Chu +2

In this paper, we provide an overview of the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) task. As large language models (LLMs) grow popular in both academia and industry, how to…

cs.CL2025

LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation

Haitao Li, Yifan Chen, Yiran Hu +7

Retrieval-augmented generation (RAG) has proven highly effective in improving large language models (LLMs) across various domains. However, there is no benchmark specifically desig…

cs.CL2025

CaseGen: A Benchmark for Multi-Stage Legal Case Documents Generation

Haitao Li, Jiaying Ye, Yiran Hu +8

Legal case documents play a critical role in judicial proceedings. As the number of cases continues to rise, the reliance on manual drafting of legal case documents is facing incre…