4 papers · 1 filter
Better Language Model Inversion by Compactly Representing Next-Token Distributions
Murtaza Nazir, Matthew Finlayson, John X. Morris +2
Language model inversion seeks to recover hidden prompts using only language model outputs. This capability has implications for security and accountability in language model deplo…
NeoBERT: A Next-Generation BERT
Lola Le Breton, Quentin Fournier, Mariam El Mezouar +2
Recent innovations in architecture, pre-training, and fine-tuning have led to the remarkable in-context learning and reasoning abilities of large auto-regressive language models su…
Nomic Embed: Training a Reproducible Long Context Text Embedder
Zach Nussbaum, John X. Morris, Brandon Duderstadt +1
This technical report describes the training of nomic-embed-text-v1, the first fully reproducible, open-source, open-weights, open-data, 8192 context length English text embedding…
DIRI: Adversarial Patient Reidentification with Large Language Models for Evaluating Clinical Text Anonymization
John X. Morris, Thomas R. Campion, Sri Laasya Nutheti +4
Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove an…