Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Thinking into the Future: Latent Lookahead Training for Transformers
Lorenzo Noci, Gregor Bachmann, Seyed-Mohsen Moosavi-Dezfooli +1
Autoregressive language models trained with next-token prediction generate text by sampling one discrete token at a time. Although very scalable, this objective forces the model to…
cs.CL2025
The pitfalls of next-token prediction
Gregor Bachmann, Vaishnavh Nagarajan
Can a mere next-token predictor faithfully model human intelligence? We crystallize this emerging concern and correct popular misconceptions surrounding it, and advocate a simple m…