1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Patrick Lutz, Aditya Gangrade, Hadi Daneshmand +1
We train a linear attention transformer on millions of masked-block matrix completion tasks: each prompt is masked low-rank matrix whose missing block may be (i) a scalar predictio…