works on

From the 1 of 18 linked papers with an AI index.

collaborators

18 papers

cs.LG2026

Disentangling the Expressivity of RoPE

Selim Jerad, Anej Svete, Jiaoda Li +1

Two accounts recur in explanations of the success of rotary position embeddings (RoPE). Expressivity studies associate periodic position information with modular predicates, wherea…

cs.LG2026

Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers

Ying Fan, Anej Svete, Kangwook Lee

The paper introduces LOTUS, a looped Transformer architecture that performs multi-step reasoning in latent space, achieving reasoning performance comparable to explicit chain-of-th…

cs.LG2026

Efficiently Representing Algorithms With Chain-of-Thought Transformers

Yanhong Li, Anej Svete, Ashish Sabharwal +1

The increasing popularity of \emph{reasoning} models -- language models that output a series of reasoning or thought tokens before producing an answer -- is justified, in part, by…

cs.LG2026

Olmo Hybrid: From Theory to Practice and Back

William Merrill, Yanhong Li, Tyler Romero +19

Recent work has demonstrated the potential of non-transformer language models, especially linear recurrent neural networks (RNNs) and hybrid models that mix recurrence and attentio…

cs.CL2026

Causally Evaluating the Learnability of Formal Language Tasks

Vésteinn Snæbjarnarson, Anej Svete, Josef Valvoda +3

Language models, as multi-task learners, acquire a wide range of abilities during training. A fundamental question is how much task-specific data is needed to learn a given task. A…

cs.LG2026

Understanding the Parameter Space Geometry of Transformers Encoding Boolean Functions

Blanka Köver, Alexandra Butoi, Anej Svete +2

Transformers consistently fail to learn certain simple functions that are provably expressible with specific parameter settings. This gap between learnability and expressivity is p…