Escaping the Big Data Paradigm with Compact Transformers
arXiv:2104.05704
Abstract
With the rise of Transformers as the standard for language processing, and their advancements in computer vision, there has been a corresponding growth in parameter size and amounts of training data. Many have come to believe that because of this, transformers are not suitable for small sets of data. This trend leads to concerns such as: limited availability of data in certain scientific domains and the exclusion of those with limited resource from research in the field. In this paper, we aim to present an approach for small-scale learning by introducing Compact Transformers. We show for the first time that with the right size, convolutional tokenization, transformers can avoid overfitting and outperform state-of-the-art CNNs on small datasets. Our models are flexible in terms of model size, and can have as little as 0.28M parameters while achieving competitive results. Our best model can reach 98% accuracy when training from scratch on CIFAR-10 with only 3.7M parameters, which is a significant improvement in data-efficiency over previous Transformer based models being over 10x smaller than other transformers and is 15% the size of ResNet50 while achieving similar performance. CCT also outperforms many modern CNN based approaches, and even some recent NAS-based approaches. Additionally, we obtain a new SOTA result on Flowers-102 with 99.76% top-1 accuracy, and improve upon the existing baseline on ImageNet (82.71% accuracy with 29% as many parameters as ViT), as well as NLP tasks. Our simple and compact design for transformers makes them more feasible to study for those with limited computing resources and/or dealing with small datasets, while extending existing research efforts in data efficient transformers. Our code and pre-trained models are publicly available at https://github.com/SHI-Labs/Compact-Transformers.
Added new results on Flowers-102, distillation
References in corpus (4)
Cited by in corpus (27)
- TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding Network
- Recent Progress in Transformer-based Medical Image Analysis
- Shuffle Transformer: Rethinking Spatial Shuffle for Vision Transformer
- An Experimental Study of Byzantine-Robust Aggregation Schemes in Federated Learning
- Improvement of Performance in Freezing of Gait detection in Parkinsons Disease using Transformer networks and a single waist worn triaxial accelerometer
- Clinically-Inspired Multi-Agent Transformers for Disease Trajectory Forecasting from Multimodal Data
- Enhancing MRI-Based Classification of Alzheimer's Disease with Explainable 3D Hybrid Compact Convolutional Transformers
- ConvMLP: Hierarchical Convolutional MLPs for Vision
- ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
- TransFusionOdom: Interpretable Transformer-based LiDAR-Inertial Fusion Odometry Estimation
- Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
- Synthesized Speech Detection Using Convolutional Transformer-Based Spectrogram Analysis
- OnDev-LCT: On-Device Lightweight Convolutional Transformers towards federated learning
- Learning Sequence Descriptor based on Spatio-Temporal Attention for Visual Place Recognition
- Vision Xformers: Efficient Attention for Image Classification
- UFO-ViT: High Performance Linear Vision Transformer without Softmax
- Activation Function Optimization Scheme for Image Classification
- Triggering Dark Showers with Conditional Dual Auto-Encoders
- Full-attention based Neural Architecture Search using Context Auto-regression
- MSN: Efficient Online Mask Selection Network for Video Instance Segmentation
- Non-Intrusive Binaural Speech Intelligibility Prediction from Discrete Latent Representations
- Harnessing Orthogonality to Train Low-Rank Neural Networks
- Hybrid BYOL-ViT: Efficient approach to deal with small datasets
- Noisy Feature Mixup
- CpT: Convolutional Point Transformer for 3D Point Cloud Processing
- Audiomer: A Convolutional Transformer For Keyword Spotting
- Contrastive Masked Autoencoders for Character-Level Open-Set Writer Identification