7 papers · 1 filter
RSGMamba: Reliability-Aware Self-Gated State Space Model for Multimodal Semantic Segmentation
Guoan Xu, Yang Xiao, Guangwei Gao +3
Multimodal semantic segmentation has emerged as a powerful paradigm for enhancing scene understanding by leveraging complementary information from multiple sensing modalities (e.g.…
Mastering Negation: Boosting Grounding Models via Grouped Opposition-Based Learning
Zesheng Yang, Xi Jiang, Bingzhang Hu +4
Current vision-language detection and grounding models predominantly focus on prompts with positive semantics and often struggle to accurately interpret and ground complex expressi…
WaveSeg: Enhancing Segmentation Precision via High-Frequency Prior and Mamba-Driven Spectrum Decomposition
Guoan Xu, Yang Xiao, Wenjing Jia +3
While recent semantic segmentation networks heavily rely on powerful pretrained encoders, most employ simplistic decoders, leading to suboptimal trade-offs between semantic context…
S2AFormer: Strip Self-Attention for Efficient Vision Transformer
Guoan Xu, Wenfeng Huang, Wenjing Jia +3
Vision Transformer (ViT) has made significant advancements in computer vision, thanks to its token mixer's sophisticated ability to capture global dependencies between all tokens.…
Cross Paradigm Representation and Alignment Transformer for Image Deraining
Shun Zou, Yi Zou, Juncheng Li +2
Transformer-based networks have achieved strong performance in low-level vision tasks like image deraining by utilizing spatial or channel-wise self-attention. However, irregular r…
Learning Dual-Domain Multi-Scale Representations for Single Image Deraining
Shun Zou, Yi Zou, Mingya Zhang +3
Existing image deraining methods typically rely on single-input, single-output, and single-scale architectures, which overlook the joint multi-scale information between external an…