20 citations · 48 across the 12 of their papers we have counts for
9 papers · 1 filter
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
Ziyang Wu, Tianjiao Ding, Yifu Lu +6
The attention operator is arguably the key distinguishing factor of transformer architectures, which have demonstrated state-of-the-art performance on a variety of tasks. However,…
A Convex Relaxation Approach to Generalization Analysis for Parallel Positively Homogeneous Networks
Uday Kiran Reddy Tadipatri, Benjamin D. Haeffele, Joshua Agterberg +1
We propose a general framework for deriving generalization bounds for parallel positively homogeneous neural networks--a class of neural networks whose input-output map decomposes…
Efficient Maximal Coding Rate Reduction by Variational Forms
Christina Baek, Ziyang Wu, Kwan Ho Ryan Chan +3
The principle of Maximal Coding Rate Reduction (MCR) has recently been proposed as a training objective for learning discriminative low-dimensional structures intrinsic to high…
Wave-Informed Matrix Factorization with Global Optimality Guarantees
Harsha Vardhan Tetali, Joel B. Harley, Benjamin D. Haeffele
With the recent success of representation learning methods, which includes deep learning as a special case, there has been considerable interest in developing representation learni…
Doubly Stochastic Subspace Clustering
Derek Lim, René Vidal, Benjamin D. Haeffele
Many state-of-the-art subspace clustering methods follow a two-step process by first constructing an affinity matrix between data points and then applying spectral clustering to th…
A Critique of Self-Expressive Deep Subspace Clustering
Benjamin D. Haeffele, Chong You, René Vidal
Subspace clustering is an unsupervised clustering technique designed to cluster data that is supported on a union of linear subspaces, with each subspace defining a cluster with di…