Don't waste SAM
arXiv:2606.10696 · doi:10.14428/esann/2023.ES2023-116
Abstract
Meta AI has recently released the Segment Anything Model (SAM), which demonstrates exceptional zero-shot image segmentation performance across various tasks with remarkable accuracy. Despite its inability to provide accurate segmentation across multiple research fields, SAM still serves as a valuable starting point for supporting the segmentation pipeline process, particularly for tasks that require extensive and senior skills annotations. This study aims to evaluate the generalization of SAM and fine-tuning SAM models using three waste segmentation datasets. Although they are captured from real scenes as SAM was pretrained on, these datasets present several challenges, including occlusions, deformable objects, transparency, and objects easily confused with backgrounds. In our findings, the fine-tuned SAM-ViT-H model outperforms the state-ofthe-art Zerowaste, and TACO datasets with a significant increase of +30 in IoU, and it closely approaches performance levels of TrashCan 1.0, with only a -1.44 difference. After evaluating these popular waste datasets, it became evident that fine-tuning SAM as a foundational model is a crucial step for providing better generalization for downstream waste segmentation tasks. Therefore, SAM should not be disregarded or wasted.
Published at European Symposium on Artificial Neural Networks (ESANN2023), Computational Intelligence and Machine Learning. Bruges (Belgium)
References in corpus (15)
- Learning Transferable Visual Models From Natural Language Supervision
- Language Models are Few-Shot Learners
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Zero-Shot Text-to-Image Generation
- TACO: Trash Annotations in Context for Litter Detection
- SAM Struggles in Concealed Scenes -- Empirical Study on Segment Anything
- Track Anything: Segment Anything Meets Videos
- Inpaint Anything: Segment Anything Meets Image Inpainting
- TrashCan: A Semantically-Segmented Dataset towards Visual Detection of Marine Debris
- Can SAM Segment Anything? When SAM Meets Camouflaged Object Detection
- SAM Fails to Segment Anything? -- SAM-Adapter: Adapting SAM in Underperformed Scenes: Camouflage, Shadow, Medical Image Segmentation, and More
- Anything-3D: Towards Single-view Anything Reconstruction in the Wild
- Learning to "Segment Anything" in Thermal Infrared Images through Knowledge Distillation with a Large Scale Dataset SATIR
- ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes
- Any-to-Any Style Transfer: Making Picasso and Da Vinci Collaborate