1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2024
Transformers on Markov Data: Constant Depth Suffices
Nived Rajaraman, Marco Bondaschi, Kannan Ramchandran +2
Attention-based transformers have been remarkably successful at modeling generative processes across various domains and modalities. In this paper, we study the behavior of transfo…
cs.LG2023★ 1 cited
Greedy Pruning with Group Lasso Provably Generalizes for Matrix Sensing
Nived Rajaraman, Devvrit, Aryan Mokhtari +1
Pruning schemes have been widely used in practice to reduce the complexity of trained models with a massive number of parameters. In fact, several practical studies have shown that…