400 citations · 403 across the 3 of their papers we have counts for
3 papers
cs.CL2024★ 3 cited
Know When To Stop: A Study of Semantic Drift in Text Generation
Ava Spataru, Eric Hambro, Elena Voita +1
In this work, we explicitly show that modern LLMs tend to generate correct facts first, then "drift away" and generate incorrect facts later: this was occasionally observed but nev…
cs.AI2024
Spectral Filters, Dark Signals, and Attention Sinks
Nicola Cancedda
Projecting intermediate representations onto the vocabulary is an increasingly popular interpretation tool for transformer-based LLMs, also known as the logit lens. We propose a qu…
cs.CL2023★ 400 cited
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì +5
Language models (LMs) exhibit remarkable abilities to solve new tasks from just a few examples or textual instructions, especially at scale. They also, paradoxically, struggle with…