Exploring Advances in Transformers and CNN for Skin Lesion Diagnosis on Small Datasets
arXiv:2205.15442 · doi:10.1007/978-3-031-21689-3_21
Abstract
Skin cancer is one of the most common types of cancer in the world. Different computer-aided diagnosis systems have been proposed to tackle skin lesion diagnosis, most of them based in deep convolutional neural networks. However, recent advances in computer vision achieved state-of-art results in many tasks, notably Transformer-based networks. We explore and evaluate advances in computer vision architectures, training methods and multimodal feature fusion for skin lesion diagnosis task. Experiments show that PiT (), CoaT () and ViT () backbone models with MetaBlock fusion achieved state-of-art results for balanced accuracy metric in PAD-UFES-20 dataset.
7 pages
References in corpus (15)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Distilling the Knowledge in a Neural Network
- A Survey on Visual Transformer
- Transformers in Vision: A Survey
- EfficientNetV2: Smaller Models and Faster Training
- Transformer in Transformer
- BEiT: BERT Pre-Training of Image Transformers
- Twins: Revisiting the Design of Spatial Attention in Vision Transformers
- Billion-scale semi-supervised learning for image classification
- Intriguing Properties of Vision Transformers
- High-Performance Large-Scale Image Recognition Without Normalization
- XCiT: Cross-Covariance Image Transformers
- iBOT: Image BERT Pre-Training with Online Tokenizer
- Deep Multimodal Fusion by Channel Exchanging
- Improved skin lesion recognition by a Self-Supervised Curricular Deep Learning approach