2 papers
cs.LG2026
Understanding the Staged Dynamics of Transformers in Learning Latent Structure
Rohan Saha, Farzane Aminmansour, Alona Fyshe
Language modeling has shown us that transformers can discover latent structure from context, but the dynamics of how they acquire different components of that structure remain poor…
cs.LG2024
Exploring Curriculum Learning for Vision-Language Tasks: A Study on Small-Scale Multimodal Training
Rohan Saha, Abrar Fahim, Alona Fyshe +1
For specialized domains, there is often not a wealth of data with which to train large machine learning models. In such limited data / compute settings, various methods exist aimin…