1 paper
Wenlin Liu, Xikun Hu, Ping Zhong
In remote sensing object detection, Convolutional Neural Networks (CNNs) excel at capturing local details while Vision Transformers (ViTs) are better at global context modeling. Ho…