7 papers
SCASeg: Strip Cross-Attention for Efficient Semantic Segmentation
Guoan Xu, Jiaming Chen, Wenfeng Huang +3
The Vision Transformer (ViT) has achieved notable success in computer vision, with its variants widely validated across various downstream tasks, including semantic segmentation. H…
Cross Paradigm Representation and Alignment Transformer for Image Deraining
Shun Zou, Yi Zou, Juncheng Li +2
Transformer-based networks have achieved strong performance in low-level vision tasks like image deraining by utilizing spatial or channel-wise self-attention. However, irregular r…
RSGMamba: Reliability-Aware Self-Gated State Space Model for Multimodal Semantic Segmentation
Guoan Xu, Yang Xiao, Guangwei Gao +3
Multimodal semantic segmentation has emerged as a powerful paradigm for enhancing scene understanding by leveraging complementary information from multiple sensing modalities (e.g.…
Mastering Negation: Boosting Grounding Models via Grouped Opposition-Based Learning
Zesheng Yang, Xi Jiang, Bingzhang Hu +4
Current vision-language detection and grounding models predominantly focus on prompts with positive semantics and often struggle to accurately interpret and ground complex expressi…
S2AFormer: Strip Self-Attention for Efficient Vision Transformer
Guoan Xu, Wenfeng Huang, Wenjing Jia +3
Vision Transformer (ViT) has made significant advancements in computer vision, thanks to its token mixer's sophisticated ability to capture global dependencies between all tokens.…
WaveSeg: Enhancing Segmentation Precision via High-Frequency Prior and Mamba-Driven Spectrum Decomposition
Guoan Xu, Yang Xiao, Wenjing Jia +3
While recent semantic segmentation networks heavily rely on powerful pretrained encoders, most employ simplistic decoders, leading to suboptimal trade-offs between semantic context…