1 citations · 2 across the 6 of their papers we have counts for
6 papers · 1 filter
Information Locality as an Inductive Bias for Neural Language Models
Taiga Someya, Anej Svete, Brian DuSell +3
Inductive biases are inherent in every machine learning system, shaping how models generalize from finite data. In the case of neural language models (LMs), debates persist as to w…
An Algorithm for Deterministic Weighted Regular Languages
Clemente Pasti, Talu Karagöz, Anej Svete +3
Extracting finite state automata (FSAs) from black-box models offers a powerful approach to gaining interpretable insights into complex model behaviors. To support this pursuit, we…
Gumbel Counterfactual Generation From Language Models
Shauli Ravfogel, Anej Svete, Vésteinn Snæbjarnarson +1
Understanding and manipulating the causal generation mechanisms in language models is essential for controlling their behavior. Previous work has primarily relied on techniques suc…
Training Neural Networks as Recognizers of Formal Languages
Alexandra Butoi, Ghazal Khalighinejad, Anej Svete +3
Characterizing the computational power of neural network architectures in terms of formal language theory remains a crucial line of research, as it describes lower and upper bounds…
Can Transformers Learn -gram Language Models?
Anej Svete, Nadav Borenstein, Mike Zhou +2
Much theoretical work has described the ability of transformers to represent formal languages. However, linking theoretical results to empirical performance is not straightforward…
The Role of -gram Smoothing in the Age of Neural Networks
Luca Malagutti, Andrius Buinovskij, Anej Svete +3
For nearly three decades, language models derived from the -gram assumption held the state of the art on the task. The key to their success lay in the application of various smo…