11 citations · 14 across the 6 of their papers we have counts for
4 papers · 1 filter
Paramixer: Parameterizing Mixing Links in Sparse Factors Works Better than Dot-Product Self-Attention
Tong Yu, Ruslan Khalitov, Lei Cheng +1
Self-Attention is a widely used building block in neural modeling to mix long-range data elements. Most self-attention neural networks employ pairwise dot-products to specify the a…
Towards Tailored Models on Private AIoT Devices: Federated Direct Neural Architecture Search
Chunhui Zhang, Xiaoming Yuan, Qianyun Zhang +3
Neural networks often encounter various stringent resource constraints while deploying on edge devices. To tackle these problems with less human efforts, automated machine learning…
Classification of Long Sequential Data using Circular Dilated Convolutional Neural Networks
Lei Cheng, Ruslan Khalitov, Tong Yu +1
Classification of long sequential data is an important Machine Learning task and appears in many application scenarios. Recurrent Neural Networks, Transformers, and Convolutional N…
Sparse Factorization of Large Square Matrices
Ruslan Khalitov, Tong Yu, Lei Cheng +1
Square matrices appear in many machine learning problems and models. Optimization over a large square matrix is expensive in memory and in time. Therefore an economic approximation…