Fair Comparison: Quantifying Variance in Resultsfor Fine-grained Visual Categorization
arXiv:2109.03156 · doi:10.1109/WACV48630.2021.00335
Abstract
For the task of image classification, researchers work arduously to develop the next state-of-the-art (SOTA) model, each bench-marking their own performance against that of their predecessors and of their peers. Unfortunately, the metric used most frequently to describe a model's performance, average categorization accuracy, is often used in isolation. As the number of classes increases, such as in fine-grained visual categorization (FGVC), the amount of information conveyed by average accuracy alone dwindles. While its most glaring weakness is its failure to describe the model's performance on a class-by-class basis, average accuracy also fails to describe how performance may vary from one trained model of the same architecture, on the same dataset, to another (both averaged across all categories and at the per-class level). We first demonstrate the magnitude of these variations across models and across class distributions based on attributes of the data, comparing results on different visual domains and different per-class image distributions, including long-tailed distributions and few-shot subsets. We then analyze the impact various FGVC methods have on overall and per-class variance. From this analysis, we both highlight the importance of reporting and comparing methods based on information beyond overall accuracy, as well as point out techniques that mitigate variance in FGVC results.
Accepted at WACV 2021; 8 pages text, 2 pages bib, 12 figures
References in corpus (14)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Neural Machine Translation by Jointly Learning to Align and Translate
- Sequence to Sequence Learning with Neural Networks
- ADADELTA: An Adaptive Learning Rate Method
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
- Deep Residual Learning for Image Recognition
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- Fine-Grained Visual Classification of Aircraft
- When Does Label Smoothing Help?
- Visualizing and Understanding Convolutional Networks
- See Better Before Looking Closer: Weakly Supervised Data Augmentation Network for Fine-Grained Visual Classification
- Lost in Translation: Loss and Decay of Linguistic Richness in Machine Translation
- Maximum-Entropy Fine-Grained Classification