activity
20232026
most citedCaseformer: Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions

4 citations · 16 across the 20 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL2026

Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories

Hexi Wang, Yujia Zhou, Bangde Du +7

Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce p…

cs.CL2026

Civil Court Simulation with Large Language Models

Yifan Chen, Haitao Li, Kaiyuan Zhang +3

Court simulation bridges legal education and judicial practice, yet human-based simulations are costly and difficult to scale. Large language models (LLMs) offer a scalable alterna…

cs.CL2026

LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks

Yifan Chen, Haitao Li, Yiran Hu +6

As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses has become essential. These tasks…

cs.CL2026★ 1 cited

LegalOne: A Family of Foundation Models for Reliable Legal Reasoning

Haitao Li, Yifan Chen, Shuo Miao +13

While Large Language Models (LLMs) have demonstrated impressive general capabilities, their direct application in the legal domain is often hindered by a lack of precise domain kno…

cs.CL2025

JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System

Weihang Su, Baoqing Yue, Qingyao Ai +6

This paper introduces JuDGE (Judgment Document Generation Evaluation), a novel benchmark for evaluating the performance of judgment document generation in the Chinese legal system.…

cs.CL2025★ 1 cited

LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation

Haitao Li, Yifan Chen, Yiran Hu +7

Retrieval-augmented generation (RAG) has proven highly effective in improving large language models (LLMs) across various domains. However, there is no benchmark specifically desig…