Leveraging Self-Supervised Vision Transformers for Segmentation-based Transfer Function Design
arXiv:2309.01408 · doi:10.1109/TVCG.2024.3401755
Abstract
In volume rendering, transfer functions are used to classify structures of interest, and to assign optical properties such as color and opacity. They are commonly defined as 1D or 2D functions that map simple features to these optical properties. As the process of designing a transfer function is typically tedious and unintuitive, several approaches have been proposed for their interactive specification. In this paper, we present a novel method to define transfer functions for volume rendering by leveraging the feature extraction capabilities of self-supervised pre-trained vision transformers. To design a transfer function, users simply select the structures of interest in a slice viewer, and our method automatically selects similar structures based on the high-level features extracted by the neural network. Contrary to previous learning-based transfer function approaches, our method does not require training of models and allows for quick inference, enabling an interactive exploration of the volume data. Our approach reduces the amount of necessary annotations by interactively informing the user about the current classification, so they can focus on annotating the structures of interest that still require annotation. In practice, this allows users to design transfer functions within seconds, instead of minutes. We compare our method to existing learning-based approaches in terms of annotation and compute time, as well as with respect to segmentation accuracy. Our accompanying video showcases the interactivity and effectiveness of our method.
accepted at TVCG 2024
References in corpus (19)
- Scikit-learn: Machine Learning in Python
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- A Simple Framework for Contrastive Learning of Visual Representations
- Improved Baselines with Momentum Contrastive Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- DINOv2: Learning Robust Visual Features without Supervision
- BEiT: BERT Pre-Training of Image Transformers
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Reproducible scaling laws for contrastive language-image learning
- Segment Anything
- iBOT: Image BERT Pre-Training with Online Tokenizer
- MISSFormer: An Effective Medical Image Segmentation Transformer
- Unsupervised Semantic Segmentation by Distilling Feature Correspondences
- Inviwo -- A Visualization System with Usage Abstraction Levels
- Optimizing Vision Transformers for Medical Image Segmentation
- Pushing the limits of self-supervised ResNets: Can we outperform supervised learning without labels on ImageNet?
- Transformer-Based Visual Segmentation: A Survey
- SAM-Med3D: Towards General-purpose Segmentation Models for Volumetric Medical Images
- MA-SAM: Modality-agnostic SAM Adaptation for 3D Medical Image Segmentation