Showing cs.LGShow all
3 papers · 1 filter
cs.LG2024
Selective Attention: Enhancing Transformer through Principled Context Control
Xuechen Zhang, Xiangyu Chang, Mingchen Li +3
The attention mechanism within the transformer architecture enables the model to weigh and combine tokens based on their relevance to the query. While self-attention has enjoyed ma…
cs.LG2024
On the Power of Convolution Augmented Transformer
Mingchen Li, Xuechen Zhang, Yixiao Huang +1
The transformer architecture has catalyzed revolutionary advances in language modeling. However, recent architectural recipes, such as state-space models, have bridged the performa…
cs.LG2024
Class-attribute Priors: Adapting Optimization to Heterogeneity and Fairness Objective
Xuechen Zhang, Mingchen Li, Jiasi Chen +2
Modern classification problems exhibit heterogeneities across individual classes: Each class may have unique attributes, such as sample size, label quality, or predictability (easy…