6 papers
Induction Heads Interpolate N-Grams
Francesco D'Angelo, Oguz Kaan Yuksel, Swathi Shree Narashiman +1
Induction heads are attention circuits believed to underlie in-context learning in transformers, yet a precise characterization of the estimators they implement remains elusive. We…
Incremental Learning of Sparse Attention Patterns in Transformers
Oğuz Kaan Yüksel, Rodrigo Alvarez Lucendo, Nicolas Flammarion
This paper studies simple transformers trained on a high-order Markov chain, where the model must incorporate information from multiple past positions, each with different statisti…
Long-Context Linear System Identification
Oğuz Kaan Yüksel, Mathieu Even, Nicolas Flammarion
This paper addresses the problem of long-context linear system identification, where the state of a dynamical system at time depends linearly on previous states ove…
Discovering Multiple and Diverse Directions for Cognitive Image Properties
Umut Kocasari, Alperen Bag, Oguz Kaan Yuksel +1
Recent research has shown that it is possible to find interpretable directions in the latent spaces of pre-trained GANs. These directions enable controllable generation and support…
Semantic Perturbations with Normalizing Flows for Improved Generalization
Oguz Kaan Yuksel, Sebastian U. Stich, Martin Jaggi +1
Data augmentation is a widely adopted technique for avoiding overfitting when training deep neural networks. However, this approach requires domain-specific knowledge and is often…
LatentCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions
Oğuz Kaan Yüksel, Enis Simsar, Ezgi Gülperi Er +1
Recent research has shown that it is possible to find interpretable directions in the latent spaces of pre-trained Generative Adversarial Networks (GANs). These directions enable c…