4 citations · 4 across the 8 of their papers we have counts for
6 papers · 1 filter
OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph
Michael J. Bommarito
We present OpenGloss, a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological his…
The KL3M Data Project: Copyright-Clean Training Resources for Large Language Models
Michael J Bommarito, Jillian Bommarito, Daniel Martin Katz
Practically all large language models have been pre-trained on data that is subject to global uncertainty related to copyright infringement and breach of contract. This creates pot…
Precise Legal Sentence Boundary Detection for Retrieval at Scale: NUPunkt and CharBoundary
Michael J Bommarito, Daniel Martin Katz, Jillian Bommarito
We present NUPunkt and CharBoundary, two sentence boundary detection libraries optimized for high-precision, high-throughput processing of legal text in large-scale applications su…
KL3M Tokenizers: A Family of Domain-Specific and Character-Level Tokenizers for Legal, Financial, and Preprocessing Applications
Michael J Bommarito, Daniel Martin Katz, Jillian Bommarito
We present the KL3M tokenizers, a family of specialized tokenizers for legal, financial, and governmental text. Despite established work on tokenization, specialized tokenizers for…
OpenEDGAR: Open Source Software for SEC EDGAR Analysis
Michael J Bommarito, Daniel Martin Katz, Eric M Detterman
OpenEDGAR is an open source Python framework designed to rapidly construct research databases based on the Electronic Data Gathering, Analysis, and Retrieval (EDGAR) system operate…
LexNLP: Natural language processing and information extraction for legal and regulatory texts
Michael J Bommarito, Daniel Martin Katz, Eric M Detterman
LexNLP is an open source Python package focused on natural language processing and machine learning for legal and regulatory text. The package includes functionality to (i) segment…