23 citations · 46 across the 14 of their papers we have counts for
9 papers · 1 filter
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
Ziyang Wu, Tianjiao Ding, Yifu Lu +6
The attention operator is arguably the key distinguishing factor of transformer architectures, which have demonstrated state-of-the-art performance on a variety of tasks. However,…
Spatial-Temporal Mixture-of-Graph-Experts for Multi-Type Crime Prediction
Ziyang Wu, Fan Liu, Jindong Han +2
As various types of crime continue to threaten public safety and economic development, predicting the occurrence of multiple types of crimes becomes increasingly vital for effectiv…
Masked Completion via Structured Diffusion with White-Box Transformers
Druv Pai, Ziyang Wu, Sam Buchanan +2
Modern learning frameworks often train deep neural networks with massive amounts of unlabeled data to learn representations by solving simple pretext tasks, then use the representa…
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
Yaodong Yu, Sam Buchanan, Druv Pai +7
In this paper, we contend that a natural objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a low-dimension…
White-Box Transformers via Sparse Rate Reduction
Yaodong Yu, Sam Buchanan, Druv Pai +5
In this paper, we contend that the objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a mixture of low-dime…
Efficient Maximal Coding Rate Reduction by Variational Forms
Christina Baek, Ziyang Wu, Kwan Ho Ryan Chan +3
The principle of Maximal Coding Rate Reduction (MCR) has recently been proposed as a training objective for learning discriminative low-dimensional structures intrinsic to high…