Publications (85)
IPDiff: Diffusion-driven ORSI Salient Object Detection with Information Reconstruction and Multi-Prior Guidance
Gongyang Li, Zhen Bai, Runmin Cong +3
Existing Salient Object Detection in Optical Remote Sensing Image (ORSI-SOD) methods mainly adopt the static inference strategy, which uses fixed trained model parameters for salie…
Stereo-GS: Multi-View Stereo Vision Model for Generalizable 3D Gaussian Splatting Reconstruction
Xiufeng Huang, Ka Chun Cheung, Runmin Cong +2
Generalizable 3D Gaussian Splatting reconstruction showcases advanced Image-to-3D content creation but requires substantial computational resources and large datasets, posing chall…
Superpixel Segmentation Based on Spatially Constrained Subspace Clustering
Hua Li, Yuheng Jia, Runmin Cong +3
Superpixel segmentation aims at dividing the input image into some representative regions containing pixels with similar and consistent intrinsic properties, without any prior know…
Query-guided Prototype Evolution Network for Few-Shot Segmentation
Runmin Cong, Hang Xiong, Jinpeng Chen +3
Previous Few-Shot Segmentation (FSS) approaches exclusively utilize support features for prototype generation, neglecting the specific requirements of the query. To address this, w…
An Underwater Image Enhancement Benchmark Dataset and Beyond
Chongyi Li, Chunle Guo, Wenqi Ren +4
Underwater image enhancement has been attracting much attention due to its significance in marine engineering and aquatic robotics. Numerous underwater image enhancement algorithms…
Feedback Chain Network For Hippocampus Segmentation
Heyu Huang, Runmin Cong, Lianhe Yang +3
The hippocampus plays a vital role in the diagnosis and treatment of many neurological disorders. Recent years, deep learning technology has made great progress in the field of med…
A Weakly Supervised Learning Framework for Salient Object Detection via Hybrid Labels
Runmin Cong, Qi Qin, Chen Zhang +4
Fully-supervised salient object detection (SOD) methods have made great progress, but such methods often rely on a large number of pixel-level annotations, which are time-consuming…
Review of Visual Saliency Detection with Comprehensive Information
Runmin Cong, Jianjun Lei, Huazhu Fu +3
Visual saliency detection model simulates the human visual system to perceive the scene, and has been widely used in many vision tasks. With the acquisition technology development,…
Saliency Detection for Stereoscopic Images Based on Depth Confidence Analysis and Multiple Cues Fusion
Runmin Cong, Jianjun Lei, Changqing Zhang +3
Stereoscopic perception is an important part of human visual system that allows the brain to perceive depth. However, depth information has not been well explored in existing salie…
DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation
Zihao Xin, Wentong Li, Yixuan Jiang +4
Vision-and-Language Navigation (VLN) requires agents to follow long-horizon instructions and navigate complex 3D environments. However, existing approaches face two major challenge…
HSCS: Hierarchical Sparsity Based Co-saliency Detection for RGBD Images
Runmin Cong, Jianjun Lei, Huazhu Fu +3
Co-saliency detection aims to discover common and salient objects in an image group containing more than two relevant images. Moreover, depth information has been demonstrated to b…
Towards Robust and Generalizable Continuous Space-Time Video Super-Resolution with Events
Shuoyan Wei, Feng Li, Shengeng Tang +4
Continuous space-time video super-resolution (C-STVSR) has garnered increasing interest for its capability to reconstruct high-resolution and high-frame-rate videos at arbitrary sp…
The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation
Hao Fang, Runmin Cong, Xiankai Lu +2
Motion expression video segmentation is designed to segment objects in accordance with the input motion expressions. In contrast to the conventional Referring Video Object Segmenta…
Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object Detection
Runmin Cong, Hongyu Liu, Chen Zhang +4
By integrating complementary information from RGB image and depth map, the ability of salient object detection (SOD) for complex and challenging scenes can be improved. In recent y…
Semantic Concentration for Self-Supervised Dense Representations Learning
Peisong Wen, Qianqian Xu, Siran Dai +2
Recent advances in image-level self-supervised learning (SSL) have made significant progress, yet learning dense representations for patches remains challenging. Mainstream methods…
GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models
Bin Wang, Ruotong Hu, Wentong Li +5
Visual and textual soft prompt tuning can effectively improve the adaptability of Vision-Language Models (VLMs) in downstream tasks. However, fine-tuning on video tasks impairs the…
SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object Detection
Jia Lin, Xiaofei Zhou, Jiyuan Liu +4
Recently segment anything model (SAM) has attracted widespread concerns, and it is often treated as a vision foundation model for universal segmentation. Some researchers have atte…
PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild
Henghui Ding, Chang Liu, Nikhila Ravi +33
This report provides a comprehensive overview of the 4th Pixel-level Video Understanding in the Wild (PVUW) Challenge, held in conjunction with CVPR 2025. It summarizes the challen…
Unleashing Correlation and Continuity for Hyperspectral Reconstruction from RGB Images
Fuxiang Feng, Runmin Cong, Shoushui Wei +4
Reconstructing Hyperspectral Images (HSI) from RGB images can yield high spatial resolution HSI at a lower cost, demonstrating significant application potential. This paper reveals…
Size-invariance Matters: Rethinking Metrics and Losses for Imbalanced Multi-object Salient Object Detection
Feiran Li, Qianqian Xu, Shilong Bao +4
This paper explores the size-invariance of evaluation metrics in Salient Object Detection (SOD), especially when multiple targets of diverse sizes co-exist in the same image. We ob…
CoADNet: Collaborative Aggregation-and-Distribution Networks for Co-Salient Object Detection
Qijian Zhang, Runmin Cong, Junhui Hou +2
Co-Salient Object Detection (CoSOD) aims at discovering salient objects that repeatedly appear in a given query group containing two or more relevant images. One challenging issue…
Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
Rui Xing, Runmin Cong, Yingying Wu +5
Understanding the dietary preferences of ancient societies and their evolution across periods and regions is crucial for revealing human-environment interactions. Seeds, as importa…
Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
Liuxin Bao, Zhihao Peng, Xiaofei Zhou +3
Medical Decision-Making (MDM) is a complex process requiring substantial domain-specific expertise to effectively synthesize heterogeneous and complicated clinical information. Whi…
Global Context-Aware Progressive Aggregation Network for Salient Object Detection
Zuyao Chen, Qianqian Xu, Runmin Cong +1
Deep convolutional neural networks have achieved competitive performance in salient object detection, in which how to learn effective and comprehensive features plays a critical ro…
Dense Attention Fluid Network for Salient Object Detection in Optical Remote Sensing Images
Qijian Zhang, Runmin Cong, Chongyi Li +5
Despite the remarkable advances in visual saliency analysis for natural scene images (NSIs), salient object detection (SOD) for optical remote sensing images (RSIs) still remains a…
BlindDiff: Empowering Degradation Modelling in Diffusion Models for Blind Image Super-Resolution
Feng Li, Yixuan Wu, Zichao Liang +4
Diffusion models (DM) have achieved remarkable promise in image super-resolution (SR). However, most of them are tailored to solving non-blind inverse problems with fixed known deg…
Divide-and-Conquer Decoupled Network for Cross-Domain Few-Shot Segmentation
Runmin Cong, Anpeng Wang, Bin Wan +3
Cross-domain few-shot segmentation (CD-FSS) aims to tackle the dual challenge of recognizing novel classes and adapting to unseen domains with limited annotations. However, encoder…
Once-for-All: Controllable Generative Image Compression with Dynamic Granularity Adaptation
Anqi Li, Feng Li, Yuxi Liu +3
Although recent generative image compression methods have demonstrated impressive potential in optimizing the rate-distortion-perception trade-off, they still face the critical cha…
PSNet: Parallel Symmetric Network for Video Salient Object Detection
Runmin Cong, Weiyu Song, Jianjun Lei +3
For the video salient object detection (VSOD) task, how to excavate the information from the appearance modality and the motion modality has always been a topic of great concern. T…
Mastering Negation: Boosting Grounding Models via Grouped Opposition-Based Learning
Zesheng Yang, Xi Jiang, Bingzhang Hu +4
Current vision-language detection and grounding models predominantly focus on prompts with positive semantics and often struggle to accurately interpret and ground complex expressi…
Co-saliency Detection for RGBD Images Based on Multi-constraint Feature Matching and Cross Label Propagation
Runmin Cong, Jianjun Lei, Huazhu Fu +3
Co-saliency detection aims at extracting the common salient regions from an image group containing two or more relevant images. It is a newly emerging topic in computer vision comm…
An Iterative Co-Saliency Framework for RGBD Images
Runmin Cong, Jianjun Lei, Huazhu Fu +4
As a newly emerging and significant topic in computer vision community, co-saliency detection aims at discovering the common salient objects in multiple related images. The existin…
BCS-Net: Boundary, Context and Semantic for Automatic COVID-19 Lung Infection Segmentation from CT Images
Runmin Cong, Haowei Yang, Qiuping Jiang +5
The spread of COVID-19 has brought a huge disaster to the world, and the automatic segmentation of infection regions can help doctors to make diagnosis quickly and reduce workload.…
Nested Network with Two-Stream Pyramid for Salient Object Detection in Optical Remote Sensing Images
Chongyi Li, Runmin Cong, Junhui Hou +3
Arising from the various object types and scales, diverse imaging orientations, and cluttered backgrounds in optical remote sensing image (RSI), it is difficult to directly extend…
Learning Detail-Structure Alternative Optimization for Blind Super-Resolution
Feng Li, Yixuan Wu, Huihui Bai +3
Existing convolutional neural networks (CNN) based image super-resolution (SR) methods have achieved impressive performance on bicubic kernel, which is not valid to handle unknown…
G2HFNet: GeoGran-Aware Hierarchical Feature Fusion Network for Salient Object Detection in Optical Remote Sensing Images
Bin Wan, Runmin Cong, Xiaofei Zhou +3
Remote sensing images captured from aerial perspectives often exhibit significant scale variations and complex backgrounds, posing challenges for salient object detection (SOD). Ex…
RGB-D Salient Object Detection with Cross-Modality Modulation and Selection
Chongyi Li, Runmin Cong, Yongri Piao +2
We present an effective method to progressively integrate and refine the cross-modality complementarities for RGB-D salient object detection (SOD). The proposed network mainly solv…
Taming Real-World Space-Time Video Super-Resolution with One-Step Diffusion
Shuoyan Wei, Feng Li, Chen Zhou +3
Diffusion models have demonstrated exceptional success in video super-resolution (VSR), exhibiting powerful capabilities for generating fine-grained details. However, their potenti…
A Parallel Down-Up Fusion Network for Salient Object Detection in Optical Remote Sensing Images
Chongyi Li, Runmin Cong, Chunle Guo +4
The diverse spatial resolutions, various object types, scales and orientations, and cluttered backgrounds in optical remote sensing images (RSIs) challenge the current salient obje…
Learning Deep Interleaved Networks with Asymmetric Co-Attention for Image Restoration
Feng Li, Runmin Cong, Huihui Bai +3
Recently, convolutional neural network (CNN) has demonstrated significant success for image restoration (IR) tasks (e.g., image super-resolution, image deblurring, rain streak remo…
Global-and-Local Collaborative Learning for Co-Salient Object Detection
Runmin Cong, Ning Yang, Chongyi Li +4
The goal of co-salient object detection (CoSOD) is to discover salient objects that commonly appear in a query group containing two or more relevant images. Therefore, how to effec…
Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline
Lingzhi He, Hongguang Zhu, Feng Li +6
Depth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR)…
DPANet: Depth Potentiality-Aware Gated Attention Network for RGB-D Salient Object Detection
Zuyao Chen, Runmin Cong, Qianqian Xu +1
There are two main issues in RGB-D salient object detection: (1) how to effectively integrate the complementarity from the cross-modal RGB-D data; (2) how to prevent the contaminat…
Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models
Zongji Yu, Wenshui Luo, Yiliu Sun +4
Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Group Relative Policy Optimizat…
NTIRE 2025 Challenge on Event-Based Image Deblurring: Methods and Results
Lei Sun, Andrea Alfarano, Peiqi Duan +85
This paper presents an overview of NTIRE 2025 the First Challenge on Event-Based Image Deblurring, detailing the proposed methodologies and corresponding results. The primary goal…
BridgeNet: A Joint Learning Network of Depth Map Super-Resolution and Monocular Depth Estimation
Qi Tang, Runmin Cong, Ronghui Sheng +4
Depth map super-resolution is a task with high practical application requirements in the industry. Existing color-guided depth map super-resolution methods usually necessitate an e…
TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning
Mingzu Liu, Hao Fang, Runmin Cong
Fine-Tuning-as-a-Service (FTaaS) facilitates the customization of Multimodal Large Language Models (MLLMs) but introduces critical backdoor risks via poisoned data. Existing defens…
Advancing Marine Research: UWSAM Framework and UIIS10K Dataset for Precise Underwater Instance Segmentation
Hua Li, Shijie Lian, Zhiyuan Li +5
With recent breakthroughs in large-scale modeling, the Segment Anything Model (SAM) has demonstrated significant potential in a variety of visual applications. However, due to the…
Cross-modality Discrepant Interaction Network for RGB-D Salient Object Detection
Chen Zhang, Runmin Cong, Qinwei Lin +4
The popularity and promotion of depth maps have brought new vigor and vitality into salient object detection (SOD), and a mass of RGB-D SOD algorithms have been proposed, mainly co…
M-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection
Jiyuan Liu, Jia Lin, Xiaofei Zhou +3
The Segment Anything Model 2 (SAM2) has emerged as a foundation model for universal segmentation. Owing to its generalizable visual representations, SAM2 has been successfully appl…
SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning
Jinpeng Chen, Runmin Cong, Yuzhi Zhao +4
Multimodal Continual Instruction Tuning (MCIT) aims to enable Multimodal Large Language Models (MLLMs) to incrementally learn new tasks without catastrophic forgetting. In this pap…
ICME 2026 Grand Challenge on Cross-Scenario Defect Detection and Fine-Grained Severity Grading for High-Precision Manufacturing
Wei Sun, Weixia Zhang, Linhan Cao +30
This paper presents the IEEE International Conference on Multimedia and Expo (ICME) 2026 Grand Challenge on Cross-Scenario Defect Detection and Fine-Grained Severity Grading for Hi…
Proto-FG3D: Prototype-based Interpretable Fine-Grained 3D Shape Classification
Shuxian Ma, Zihao Dong, Runmin Cong +2
Deep learning-based multi-view coarse-grained 3D shape classification has achieved remarkable success over the past decade, leveraging the powerful feature learning capabilities of…
From Sight to Insight: Unleashing Eye-Tracking in Weakly Supervised Video Salient Object Detection
Qi Qin, Runmin Cong, Gen Zhan +2
The eye-tracking video saliency prediction (VSP) task and video salient object detection (VSOD) task both focus on the most attractive objects in video and show the result in the f…
Beyond Global Scanning: Adaptive Visual State Space Modeling for Salient Object Detection in Optical Remote Sensing Images
Mengyu Ren, Yutong Li, Hua Li +2
Salient object detection (SOD) in optical remote sensing images (ORSIs) faces numerous challenges, including significant variations in target scales and low contrast between target…
UNINEXT-Cutie: The 1st Solution for LSVOS Challenge RVOS Track
Hao Fang, Feiyu Pan, Xiankai Lu +2
Referring video object segmentation (RVOS) relies on natural language expressions to segment target objects in video. In this year, LSVOS Challenge RVOS Track replaced the origin Y…
Strike a Balance in Continual Panoptic Segmentation
Jinpeng Chen, Runmin Cong, Yuxuan Luo +2
This study explores the emerging area of continual panoptic segmentation, highlighting three key balances. First, we introduce past-class backtrace distillation to balance the stab…
AIM 2024 Challenge on Video Saliency Prediction: Methods and Results
Andrey Moskalenko, Alexey Bryncev, Dmitry Vatolin +30
This paper reviews the Challenge on Video Saliency Prediction at AIM 2024. The goal of the participants was to develop a method for predicting accurate saliency maps for the provid…
Enhanced Quality Aware-Scalable Underwater Image Compression
Linwei Zhu, Junhao Zhu, Xu Zhang +4
Underwater imaging plays a pivotal role in marine exploration and ecological monitoring. However, it faces significant challenges of limited transmission bandwidth and severe disto…
CIR-Net: Cross-modality Interaction and Refinement for RGB-D Salient Object Detection
Runmin Cong, Qinwei Lin, Chen Zhang +4
Focusing on the issue of how to effectively capture and utilize cross-modality information in RGB-D salient object detection (SOD) task, we present a convolutional neural network (…
UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken
Runmin Cong, Zongji Yu, Hao Fang +2
Underwater Instance Segmentation (UIS) tasks are crucial for underwater complex scene detection. Mamba, as an emerging state space model with inherently linear complexity and globa…
Boundary Guided Semantic Learning for Real-time COVID-19 Lung Infection Segmentation System
Runmin Cong, Yumo Zhang, Ning Yang +6
The coronavirus disease 2019 (COVID-19) continues to have a negative impact on healthcare systems around the world, though the vaccines have been developed and national vaccination…
Underwater Image Enhancement via Medium Transmission-Guided Multi-Color Space Embedding
Chongyi Li, Saeed Anwar, Junhui Hou +3
Underwater images suffer from color casts and low contrast due to wavelength- and distance-dependent attenuation and scattering. To solve these two degradation issues, we present a…
Learning Hierarchical Color Guidance for Depth Map Super-Resolution
Runmin Cong, Ronghui Sheng, Hao Wu +5
Color information is the most commonly used prior knowledge for depth map super-resolution (DSR), which can provide high-frequency boundary guidance for detail restoration. However…
SDDNet: Style-guided Dual-layer Disentanglement Network for Shadow Detection
Runmin Cong, Yuchen Guan, Jinpeng Chen +3
Despite significant progress in shadow detection, current methods still struggle with the adverse impact of background color, which may lead to errors when shadows are present on c…
Multi-Projection Fusion and Refinement Network for Salient Object Detection in 360° Omnidirectional Image
Runmin Cong, Ke Huang, Jianjun Lei +3
Salient object detection (SOD) aims to determine the most visually attractive objects in an image. With the development of virtual reality technology, 360° omnidirectional image h…
Exploring Resolution Fields for Scalable Image Compression with Uncertainty Guidance
Dongyi Zhang, Feng Li, Man Liu +4
Recently, there are significant advancements in learning-based image compression methods surpassing traditional coding standards. Most of them prioritize achieving the best rate-di…
LSVOS Challenge Report: Large-scale Complex and Long Video Object Segmentation
Henghui Ding, Lingyi Hong, Chang Liu +30
Despite the promising performance of current video segmentation models on existing benchmarks, these models still struggle with complex scenes. In this paper, we introduce the 6th…
Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline
Tiantian Geng, Teng Wang, Jinming Duan +2
Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos o…
RSONet: Region-guided Selective Optimization Network for RGB-T Salient Object Detection
Bin Wan, Runmin Cong, Xiaofei Zhou +3
This paper focuses on the inconsistency in salient regions between RGB and thermal images. To address this issue, we propose the Region-guided Selective Optimization Network for RG…
Rethinking Conditional Generation for Underwater Salient Object Detection
Hua Li, Yongjie Weng, Yutong Li +3
Salient Object Detection in underwater images remains challenging due to low contrast, uneven illumination, and color distortion caused by scattering and absorption effects, which…
Does Thermal Really Always Matter for RGB-T Salient Object Detection?
Runmin Cong, Kepu Zhang, Chen Zhang +4
In recent years, RGB-T salient object detection (SOD) has attracted continuous attention, which makes it possible to identify salient objects in environments such as low light by i…
Stereo Superpixel Segmentation Via Decoupled Dynamic Spatial-Embedding Fusion Network
Hua Li, Junyan Liang, Ruiqi Wu +3
Stereo superpixel segmentation aims at grouping the discretizing pixels into perceptual regions through left and right views more collaboratively and efficiently. Existing superpix…
Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement
Chunle Guo, Chongyi Li, Jichang Guo +4
The paper presents a novel method, Zero-Reference Deep Curve Estimation (Zero-DCE), which formulates light enhancement as a task of image-specific curve estimation with a deep netw…
RDNet: Region Proportion-Aware Dynamic Adaptive Salient Object Detection Network in Optical Remote Sensing Images
Bin Wan, Runmin Cong, Xiaofei Zhou +3
Salient object detection (SOD) in remote sensing images faces significant challenges due to large variations in object sizes, the computational cost of self-attention mechanisms, a…
Frequency Perception Network for Camouflaged Object Detection
Runmin Cong, Mengyao Sun, Sanyi Zhang +3
Camouflaged object detection (COD) aims to accurately detect objects hidden in the surrounding environment. However, the existing COD methods mainly locate camouflaged objects in t…
NuI-Go: Recursive Non-Local Encoder-Decoder Network for Retinal Image Non-Uniform Illumination Removal
Chongyi Li, Huazhu Fu, Runmin Cong +2
Retinal images have been widely used by clinicians for early diagnosis of ocular diseases. However, the quality of retinal images is often clinically unsatisfactory due to eye lesi…
Bridging Component Learning with Degradation Modelling for Blind Image Super-Resolution
Yixuan Wu, Feng Li, Huihui Bai +3
Convolutional Neural Network (CNN)-based image super-resolution (SR) has exhibited impressive success on known degraded low-resolution (LR) images. However, this type of approach i…
You Can Mask More For Extremely Low-Bitrate Image Compression
Anqi Li, Feng Li, Jiaxin Han +6
Learned image compression (LIC) methods have experienced significant progress during recent years. However, these methods are primarily dedicated to optimizing the rate-distortion…
TDS-CLIP: Temporal Difference Side Network for Efficient VideoAction Recognition
Bin Wang, Wentong Li, Wenqian Wang +3
Recently, large-scale pre-trained vision-language models (e.g., CLIP), have garnered significant attention thanks to their powerful representative capabilities. This inspires resea…
RRNet: Relational Reasoning Network with Parallel Multi-scale Attention for Salient Object Detection in Optical Remote Sensing Images
Runmin Cong, Yumo Zhang, Leyuan Fang +3
Salient object detection (SOD) for optical remote sensing images (RSIs) aims at locating and extracting visually distinctive objects/regions from the optical RSIs. Despite some sal…
Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter
Zhiyang Chen, Chen Zhang, Hao Fang +1
Underwater instance segmentation (UIS), integrating pixel-level understanding and instance-level discrimination, is a pivotal technology in marine resource exploration and ecologic…
PUGAN: Physical Model-Guided Underwater Image Enhancement Using GAN with Dual-Discriminators
Runmin Cong, Wenyu Yang, Wei Zhang +4
Due to the light absorption and scattering induced by the water medium, underwater images usually suffer from some degradation problems, such as low contrast, color distortion, and…
Video Object Segmentation via SAM 2: The 4th Solution for LSVOS Challenge VOS Track
Feiyu Pan, Hao Fang, Runmin Cong +2
Video Object Segmentation (VOS) task aims to segmenting a particular object instance throughout the entire video sequence given only the object mask of the first frame. Recently, S…
Diving into Underwater: Segment Anything Model Guided Underwater Salient Instance Segmentation and A Large-scale Dataset
Shijie Lian, Ziyi Zhang, Hua Li +4
With the breakthrough of large models, Segment Anything Model (SAM) and its extensions have been attempted to apply in diverse tasks of computer vision. Underwater salient instance…