6 papers
Learning to Reason in 13 Parameters
John X. Morris, Niloofar Mireshghallah, Mark Ibrahim +1
Recent research has shown that language models can learn to \textit{reason}, often via reinforcement learning. Some work even trains low-rank parameterizations for reasoning, but c…
Better Language Model Inversion by Compactly Representing Next-Token Distributions
Murtaza Nazir, Matthew Finlayson, John X. Morris +2
Language model inversion seeks to recover hidden prompts using only language model outputs. This capability has implications for security and accountability in language model deplo…
NeoBERT: A Next-Generation BERT
Lola Le Breton, Quentin Fournier, Mariam El Mezouar +2
Recent innovations in architecture, pre-training, and fine-tuning have led to the remarkable in-context learning and reasoning abilities of large auto-regressive language models su…
Nomic Embed: Training a Reproducible Long Context Text Embedder
Zach Nussbaum, John X. Morris, Brandon Duderstadt +1
This technical report describes the training of nomic-embed-text-v1, the first fully reproducible, open-source, open-weights, open-data, 8192 context length English text embedding…
Corpus Poisoning via Approximate Greedy Gradient Descent
Jinyan Su, Preslav Nakov, Claire Cardie
Dense retrievers are widely used in information retrieval and have also been successfully extended to other knowledge intensive areas such as language models, e.g., Retrieval-Augme…
DIRI: Adversarial Patient Reidentification with Large Language Models for Evaluating Clinical Text Anonymization
John X. Morris, Thomas R. Campion, Sri Laasya Nutheti +4
Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove an…