1 paper
Somesh Mehra, Javier Alonso Garcia, Lukas Mauch
We systematically investigate multi-token prediction (MTP) capabilities within LLMs pre-trained for next-token prediction (NTP). We first show that such models inherently possess M…