activity
20242026
collaborators

15 papers

cs.CL2026

DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation

Li Zhang, Yuzhen Shi, Yiran Hu +15

Lawyer-client consultation is a critical starting point for legal services. Effective legal assistance hinges on eliciting sufficient and truthful information from clients in order…

cs.AI2026

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +306

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…

cs.CL2026

LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks

Yifan Chen, Haitao Li, Yiran Hu +6

As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses has become essential. These tasks…

cs.CY2026

AppellateGen: A Benchmark for Appellate Legal Judgment Generation

Hongkun Yang, Lionel Z. Wang, Wei Fan +10

Legal judgment generation is a critical task in legal intelligence. However, existing research in legal judgment generation has predominantly focused on first-instance trials, rely…

cs.CL2026

PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice

Yuzhen Shi, Huanghai Liu, Yiran Hu +27

As large language models (LLMs) are increasingly applied to legal domain-specific tasks, evaluating their ability to perform legal work in real-world settings has become essential.…

cs.CY2026

Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions

Yiran Hu, Huanghai Liu, Chong Wang +15

Large language models (LLMs) are being increasingly integrated into legal applications, including judicial decision support, legal practice assistance, and public-facing legal serv…