1 citations · 1 across the 2 of their papers we have counts for
11 papers
Gemma 4 Technical Report
Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemm…
Rotary Position Encodings for Graphs
Isaac Reid, Arijit Sehanobish, Cederik Höfs +7
We study the extent to which rotary position encodings (RoPE), a recent transformer position encoding algorithm broadly adopted in large language models (LLMs) and vision transform…
The Illusion of Stochasticity in LLMs
Xiangming Gu, Soham De, Michalis Titsias +3
In this work, we demonstrate that reliable stochastic sampling is a fundamental yet unfulfilled requirement for Large Language Models (LLMs) operating as agents. Agentic systems ar…
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
Xiangming Gu, Soham De, Larisa Markeeva +2
Large Reasoning Models (LRMs) have shown remarkable performance on challenging questions, such as math and coding. However, to obtain a high quality solution, one may need to sampl…
Mining Generalizable Activation Functions
Alex Vitvitskyi, Michael Boratko, Matej Grcic +3
The choice of activation function is an active area of research, with different proposals aimed at improving optimization, while maintaining expressivity. Additionally, the activat…
Perplexity Cannot Always Tell Right from Wrong
Petar VeliÄkoviÄ, Federico Barbero, Christos Perivolaropoulos +2
Perplexity -- a function measuring a model's overall level of "surprise" when encountering a particular output -- has gained significant traction in recent years, both as a loss fu…