Showing cs.LGShow all
3 papers · 1 filter
cs.LG2024
Transformers Can Navigate Mazes With Multi-Step Prediction
Niklas Nolte, Ouail Kitouni, Adina Williams +2
Despite their remarkable success in language modeling, transformers trained to predict the next token in a sequence struggle with long-term planning. This limitation is particularl…
cs.LG2024
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
Ouail Kitouni, Niklas Nolte, Diane Bouchacourt +3
Today's best language models still struggle with hallucinations: factually incorrect generations, which impede their ability to reliably retrieve information seen during training.…
cs.LG2024
DP-RDM: Adapting Diffusion Models to Private Domains Without Fine-Tuning
Jonathan Lebensold, Maziar Sanjabi, Pietro Astolfi +4
Text-to-image diffusion models have been shown to suffer from sample-level memorization, possibly reproducing near-perfect replica of images that they are trained on, which may be…