BERT Goes to Law School: Quantifying the Competitive Advantage of Access to Large Legal Corpora in Contract Understanding
arXiv:1911.00473
Abstract
Fine-tuning language models, such as BERT, on domain specific corpora has proven to be valuable in domains like scientific papers and biomedical text. In this paper, we show that fine-tuning BERT on legal documents similarly provides valuable improvements on NLP tasks in the legal domain. Demonstrating this outcome is significant for analyzing commercial agreements, because obtaining large legal corpora is challenging due to their confidential nature. As such, we show that having access to large legal corpora is a competitive advantage for commercial applications, and academic research on analyzing contracts.
Cited by in corpus (9)
- On the Opportunities and Risks of Foundation Models
- CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From Characters
- Lawformer: A Pre-trained Language Model for Chinese Legal Long Documents
- When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset
- Privacy-Preserving Models for Legal Natural Language Processing
- To Tune or Not To Tune? Zero-shot Models for Legal Case Entailment
- A Small Claims Court for the NLP: Judging Legal Text Classification Strategies With Small Datasets
- JuriBERT: A Masked-Language Model Adaptation for French Legal Text
- Towards Grad-CAM Based Explainability in a Legal Text Processing Pipeline