7 citations · 7 across the 5 of their papers we have counts for
5 papers · 1 filter
Natural Language Processing in the Legal Domain
Dirk Hartung, Daniel Martin Katz, Michael J. Bommarito +3
We summarize the current state of the field of NLP & Law with a specific focus on recent technical and substantive developments. To support our analysis, we construct and analyze a…
OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph
Michael J. Bommarito
We present OpenGloss, a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological his…
The KL3M Data Project: Copyright-Clean Training Resources for Large Language Models
Michael J Bommarito, Jillian Bommarito, Daniel Martin Katz
Practically all large language models have been pre-trained on data that is subject to global uncertainty related to copyright infringement and breach of contract. This creates pot…
Precise Legal Sentence Boundary Detection for Retrieval at Scale: NUPunkt and CharBoundary
Michael J Bommarito, Daniel Martin Katz, Jillian Bommarito
We present NUPunkt and CharBoundary, two sentence boundary detection libraries optimized for high-precision, high-throughput processing of legal text in large-scale applications su…
KL3M Tokenizers: A Family of Domain-Specific and Character-Level Tokenizers for Legal, Financial, and Preprocessing Applications
Michael J Bommarito, Daniel Martin Katz, Jillian Bommarito
We present the KL3M tokenizers, a family of specialized tokenizers for legal, financial, and governmental text. Despite established work on tokenization, specialized tokenizers for…