10 citations · 11 across the 8 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation
Yuxin Ren, Maxwell D Collins, Miao Hu +1
Self-attention serves as the core foundation of large-scale transformer pretraining, but its quadratic token interaction cost makes inference expensive. Replacing attention with si…
cs.LG2022
Tackling Instance-Dependent Label Noise with Dynamic Distribution Calibration
Manyi Zhang, Yuxin Ren, Zihao Wang +1
Instance-dependent label noise is realistic but rather challenging, where the label-corruption process depends on instances directly. It causes a severe distribution shift between…