Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
arXiv:2403.18593 · doi:10.1016/j.isprsjprs.2024.09.009
Abstract
The tokenizer, as one of the fundamental components of large models, has long been overlooked or even misunderstood in visual tasks. One key factor of the great comprehension power of the large language model is that natural language tokenizers utilize meaningful words or subwords as the basic elements of language. In contrast, mainstream visual tokenizers, represented by patch-based methods such as Patch Embed, rely on meaningless rectangular patches as basic elements of vision, which cannot serve as effectively as words or subwords in language. Starting from the essence of the tokenizer, we defined semantically independent regions (SIRs) for vision. We designed a simple HOmogeneous visual tOKenizer: HOOK. HOOK mainly consists of two modules: the Object Perception Module (OPM) and the Object Vectorization Module (OVM). To achieve homogeneity, the OPM splits the image into 4*4 pixel seeds and then utilizes the attention mechanism to perceive SIRs. The OVM employs cross-attention to merge seeds within the same SIR. To achieve adaptability, the OVM defines a variable number of learnable vectors as cross-attention queries, allowing for the adjustment of token quantity. We conducted experiments on the NWPU-RESISC45, WHU-RS19 classification dataset, and GID5 segmentation dataset for sparse and dense tasks. The results demonstrate that the visual tokens obtained by HOOK correspond to individual objects, which demonstrates homogeneity. HOOK outperformed Patch Embed by 6\% and 10\% in the two tasks and achieved state-of-the-art performance compared to the baselines used for comparison. Compared to Patch Embed, which requires more than one hundred tokens for one image, HOOK requires only 6 and 8 tokens for sparse and dense tasks, respectively, resulting in efficiency improvements of 1.5 to 2.8 times. The code is available at https://github.com/GeoX-Lab/Hook.
24 pages, 9 figures, 8 tables
References in corpus (25)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Efficient Estimation of Word Representations in Vector Space
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- LLaMA: Open and Efficient Foundation Language Models
- Remote Sensing Image Scene Classification: Benchmark and State of the Art
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Longformer: The Long-Document Transformer
- PaLM: Scaling Language Modeling with Pathways
- Visual Instruction Tuning
- Visual Transformers: Token-based Image Representation and Processing for Computer Vision
- Early Convolutions Help Transformers See Better
- DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Vision Transformer for Small-Size Datasets
- CMID: A Unified Self-Supervised Learning Framework for Remote Sensing Image Understanding
- Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision Transformers
- Vision Transformers with Patch Diversification
- AdaViT: Adaptive Tokens for Efficient Vision Transformer
- Human-Object Interaction Detection:A Quick Survey and Examination of Methods
- Extending global-local view alignment for self-supervised learning with remote sensing imagery
- TOV: The Original Vision Model for Optical Remote Sensing Image Understanding via Self-supervised Learning
- Visual Concepts Tokenization
- SPFormer: Enhancing Vision Transformer with Superpixel Representation
- SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model