2 papers
cs.CV2024
Centroid-centered Modeling for Efficient Vision Transformer Pre-training
Xin Yan, Zuchao Li, Lefei Zhang
Masked Image Modeling (MIM) is a new self-supervised vision pre-training paradigm using a Vision Transformer (ViT). Previous works can be pixel-based or token-based, using original…
cs.CV2024
Expediting Building Footprint Extraction from High-resolution Remote Sensing Images via progressive lenient supervision
Haonan Guo, Bo Du, Chen Wu +2
The efficacy of building footprint segmentation from remotely sensed images has been hindered by model transfer effectiveness. Many existing building segmentation methods were deve…