◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Jillian Bommarito

4 papers hereh-index 488 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2
  • last author2

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL3
  • cs.CY1

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.CL2025

The KL3M Data Project: Copyright-Clean Training Resources for Large Language Models

Michael J Bommarito, Jillian Bommarito, Daniel Martin Katz

Practically all large language models have been pre-trained on data that is subject to global uncertainty related to copyright infringement and breach of contract. This creates pot…

cs.CL2025

Precise Legal Sentence Boundary Detection for Retrieval at Scale: NUPunkt and CharBoundary

Michael J Bommarito, Daniel Martin Katz, Jillian Bommarito

We present NUPunkt and CharBoundary, two sentence boundary detection libraries optimized for high-precision, high-throughput processing of legal text in large-scale applications su…

cs.CL2025

KL3M Tokenizers: A Family of Domain-Specific and Character-Level Tokenizers for Legal, Financial, and Preprocessing Applications

Michael J Bommarito, Daniel Martin Katz, Jillian Bommarito

We present the KL3M tokenizers, a family of specialized tokenizers for legal, financial, and governmental text. Despite established work on tokenization, specialized tokenizers for…

cs.CY2025

Towards Best Practices for Open Datasets for LLM Training

Stefan Baack, Stella Biderman, Kasia Odrozek +36

Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.