1 paper · 1 filter
Adam S. Shai, Sarah E. Marzen, Lucas Teixeira +2
What computational structure are we building into large language models when we train them on next-token prediction? Here, we present evidence that this structure is given by the m…