activity
20182022
most citedCAIL2019-SCM: A Dataset of Similar Case Matching in Legal Domain

35 citations · 108 across the 9 of their papers we have counts for

collaborators

15 papers

cs.LG2022

A Roadmap for Big Model

Sha Yuan, Hanyu Zhao, Shuai Zhao +97

With the rapid development of deep learning, training Big Models (BMs) for multiple downstream tasks becomes a popular paradigm. Researchers have achieved various outcomes in the c…

cs.CL2022

LEVEN: A Large-Scale Chinese Legal Event Detection Dataset

Feng Yao, Chaojun Xiao, Xiaozhi Wang +7

Recognizing facts is the most fundamental step in making judgments, hence detecting events in the legal documents is important to legal case analysis tasks. However, existing Legal…

cs.CL202115 cited

CPM-2: Large-scale Cost-effective Pre-trained Language Models

Zhengyan Zhang, Yuxian Gu, Xu Han +16

In recent years, the size of pre-trained language models (PLMs) has grown by leaps and bounds. However, efficiency issues of these large-scale PLMs limit their utilization in real-…

cs.CL2021

Lawformer: A Pre-trained Language Model for Chinese Legal Long Documents

Chaojun Xiao, Xueyu Hu, Zhiyuan Liu +2

Legal artificial intelligence (LegalAI) aims to benefit legal systems with the technology of artificial intelligence, especially natural language processing (NLP). Recently, inspir…

cs.CL202111 cited

Equality before the Law: Legal Judgment Consistency Analysis for Fairness

Yuzhong Wang, Chaojun Xiao, Shirong Ma +5

In a legal system, judgment consistency is regarded as one of the most important manifestations of fairness. However, due to the complexity of factual elements that impact sentenci…

cs.IR202114 cited

UPRec: User-Aware Pre-training for Recommender Systems

Chaojun Xiao, Ruobing Xie, Yuan Yao +4

Existing sequential recommendation methods rely on large amounts of training data and usually suffer from the data sparsity problem. To tackle this, the pre-training mechanism has…