most citedImproving underwater semantic segmentation with underwater image quality attention and muti-scale aggregation attention

9 citations · 9 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2025

SFFR: Spatial-Frequency Feature Reconstruction for Multispectral Aerial Object Detection

Xin Zuo, Chenyu Qu, Haibo Zhan +2

Recent multispectral object detection methods have primarily focused on spatial-domain feature fusion based on CNNs or Transformers, while the potential of frequency-domain feature…

cs.CV2025

IRDFusion: Iterative Relation-Map Difference guided Feature Fusion for Multispectral Object Detection

Jifeng Shen, Haibo Zhan, Xin Zuo +4

Current multispectral object detection methods often retain extraneous background or noise during feature fusion, limiting perceptual performance. To address this, we propose an in…

cs.CV2025

InceptionMamba: An Efficient Hybrid Network with Large Band Convolution and Bottleneck Mamba

Yuhang Wang, Jun Li, Zhijian Wu +3

Within the family of convolutional neural networks, InceptionNeXt has shown excellent competitiveness in image classification and a number of downstream tasks. Built on parallel on…

cs.CV2025

Multispectral State-Space Feature Fusion: Bridging Shared and Cross-Parametric Interactions for Object Detection

Jifeng Shen, Haibo Zhan, Shaohua Dong +3

Modern multispectral feature fusion for object detection faces two critical limitations: (1) Excessive preference for local complementary features over cross-modal shared semantics…

cs.CV2025★ 9 cited

Improving underwater semantic segmentation with underwater image quality attention and muti-scale aggregation attention

Xin Zuo, Jiaran Jiang, Jifeng Shen +1

Underwater image understanding is crucial for both submarine navigation and seabed exploration. However, the low illumination in underwater environments degrades the imaging qualit…

cs.CV2025

Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints

Ming Dai, Jian Li, Jiedong Zhuang +2

Multi-task visual grounding involves the simultaneous execution of localization and segmentation in images based on textual expressions. The majority of advanced methods predominan…