1 citations · 1 across the 2 of their papers we have counts for
9 papers
Gemma 4 Technical Report
Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemm…
Rotary Position Encodings for Graphs
Isaac Reid, Arijit Sehanobish, Cederik Höfs +7
We study the extent to which rotary position encodings (RoPE), a recent transformer position encoding algorithm broadly adopted in large language models (LLMs) and vision transform…
How do LLMs Compute Verbal Confidence
Dharshan Kumaran, Arthur Conmy, Federico Barbero +3
Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from black-box models. However, how LLMs in…
Perplexity Cannot Always Tell Right from Wrong
Petar VeliÄkoviÄ, Federico Barbero, Christos Perivolaropoulos +2
Perplexity -- a function measuring a model's overall level of "surprise" when encountering a particular output -- has gained significant traction in recent years, both as a loss fu…
Extracting alignment data in open models
Federico Barbero, Xiangming Gu, Christopher A. Choquette-Choo +6
In this work, we show that it is possible to extract significant amounts of alignment training data from a post-trained model -- useful to steer the model to improve certain capabi…
Why do LLMs attend to the first token?
Federico Barbero, Ãlvaro Arroyo, Xiangming Gu +4
Large Language Models (LLMs) tend to attend heavily to the first token in the sequence -- creating a so-called attention sink. Many works have studied this phenomenon in detail, pr…