5 citations · 5 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
DynMuon: A Dynamic Spectral Shaping View of Muon
Fangzhou Wu, Rikhav Shah, Sandeep Silwal +1
In recent years, Muon has emerged as the dominant method for training large language models, and transformers more broadly. The essential difference, when compared to standard grad…
cs.LG2019★ 5 cited
Using Dimensionality Reduction to Optimize t-SNE
Rikhav Shah, Sandeep Silwal
t-SNE is a popular tool for embedding multi-dimensional datasets into two or three dimensions. However, it has a large computational cost, especially when the input data has many d…