6 citations · 7 across the 13 of their papers we have counts for
14 papers
Super Apriel: One Checkpoint, Many Speeds
SLAM Labs, :, Oleksiy Ostapenko +13
We release Super Apriel, a 15B-parameter supernet in which every decoder layer provides four trained mixer choices -- Full Attention (FA), Sliding Window Attention (SWA), Kimi Delt…
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
Alireza Mousavi-Hosseini, Murat A. Erdogdu
We study post-training linear autoregressive models with outcome and process rewards. Given a context , the model must predict the response …
From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGD
Konstantinos Christopher Tsiolis, Alireza Mousavi-Hosseini, Murat A. Erdogdu
To understand feature learning dynamics in neural networks, recent theoretical works have focused on gradient-based learning of Gaussian single-index models, where the label is a n…
Flow Matching with Semidiscrete Couplings
Alireza Mousavi-Hosseini, Stephen Y. Zhang, Michal Klein +1
Flow models parameterized as time-dependent velocity fields can generate data from noise by integrating an ODE. These models are often trained using flow matching, i.e. by sampling…
On Fitting Flow Models with Large Sinkhorn Couplings
Stephen Zhang, Alireza Mousavi-Hosseini, Michal Klein +1
Flow models transform data gradually from one modality (e.g. noise) onto another (e.g. images). Such models are parameterized by a time-dependent velocity field, trained to fit seg…
When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective
Alireza Mousavi-Hosseini, Clayton Sanford, Denny Wu +1
Theoretical efforts to prove advantages of Transformers in comparison with classical architectures such as feedforward and recurrent neural networks have mostly focused on represen…