BERT Goes to Law School: Quantifying the Competitive Advantage of Access to Large Legal Corpora in Contract Understanding
arXiv:1911.00473
Abstract
Fine-tuning language models, such as BERT, on domain specific corpora has proven to be valuable in domains like scientific papers and biomedical text. In this paper, we show that fine-tuning BERT on legal documents similarly provides valuable improvements on NLP tasks in the legal domain. Demonstrating this outcome is significant for analyzing commercial agreements, because obtaining large legal corpora is challenging due to their confidential nature. As such, we show that having access to large legal corpora is a competitive advantage for commercial applications, and academic research on analyzing contracts.
Cited by in corpus (6)
- CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From Characters
- To Tune or Not To Tune? Zero-shot Models for Legal Case Entailment
- Privacy-Preserving Models for Legal Natural Language Processing
- A Small Claims Court for the NLP: Judging Legal Text Classification Strategies With Small Datasets
- JuriBERT: A Masked-Language Model Adaptation for French Legal Text
- Towards Grad-CAM Based Explainability in a Legal Text Processing Pipeline