activity
20242026
collaborators

6 papers

cs.LG2026

Learning to Reason in 13 Parameters

John X. Morris, Niloofar Mireshghallah, Mark Ibrahim +1

Recent research has shown that language models can learn to \textit{reason}, often via reinforcement learning. Some work even trains low-rank parameterizations for reasoning, but c…

cs.CL2025

Better Language Model Inversion by Compactly Representing Next-Token Distributions

Murtaza Nazir, Matthew Finlayson, John X. Morris +2

Language model inversion seeks to recover hidden prompts using only language model outputs. This capability has implications for security and accountability in language model deplo…

cs.CL2025

NeoBERT: A Next-Generation BERT

Lola Le Breton, Quentin Fournier, Mariam El Mezouar +2

Recent innovations in architecture, pre-training, and fine-tuning have led to the remarkable in-context learning and reasoning abilities of large auto-regressive language models su…

cs.CL2025

Nomic Embed: Training a Reproducible Long Context Text Embedder

Zach Nussbaum, John X. Morris, Brandon Duderstadt +1

This technical report describes the training of nomic-embed-text-v1, the first fully reproducible, open-source, open-weights, open-data, 8192 context length English text embedding…

cs.IR2024

Corpus Poisoning via Approximate Greedy Gradient Descent

Jinyan Su, Preslav Nakov, Claire Cardie

Dense retrievers are widely used in information retrieval and have also been successfully extended to other knowledge intensive areas such as language models, e.g., Retrieval-Augme…

cs.CL2024

DIRI: Adversarial Patient Reidentification with Large Language Models for Evaluating Clinical Text Anonymization

John X. Morris, Thomas R. Campion, Sri Laasya Nutheti +4

Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove an…