papers

Publications (85)

cs.CV2026

IPDiff: Diffusion-driven ORSI Salient Object Detection with Information Reconstruction and Multi-Prior Guidance

Gongyang Li, Zhen Bai, Runmin Cong +3

Existing Salient Object Detection in Optical Remote Sensing Image (ORSI-SOD) methods mainly adopt the static inference strategy, which uses fixed trained model parameters for salie…

cs.CV2026

Stereo-GS: Multi-View Stereo Vision Model for Generalizable 3D Gaussian Splatting Reconstruction

Xiufeng Huang, Ka Chun Cheung, Runmin Cong +2

Generalizable 3D Gaussian Splatting reconstruction showcases advanced Image-to-3D content creation but requires substantial computational resources and large datasets, posing chall…

cs.CV2020

Superpixel Segmentation Based on Spatially Constrained Subspace Clustering

Hua Li, Yuheng Jia, Runmin Cong +3

Superpixel segmentation aims at dividing the input image into some representative regions containing pixels with similar and consistent intrinsic properties, without any prior know…

cs.CV2024

Query-guided Prototype Evolution Network for Few-Shot Segmentation

Runmin Cong, Hang Xiong, Jinpeng Chen +3

Previous Few-Shot Segmentation (FSS) approaches exclusively utilize support features for prototype generation, neglecting the specific requirements of the query. To address this, w…

cs.CV2019

An Underwater Image Enhancement Benchmark Dataset and Beyond

Chongyi Li, Chunle Guo, Wenqi Ren +4

Underwater image enhancement has been attracting much attention due to its significance in marine engineering and aquatic robotics. Numerous underwater image enhancement algorithms…

cs.CV2022

Feedback Chain Network For Hippocampus Segmentation

Heyu Huang, Runmin Cong, Lianhe Yang +3

The hippocampus plays a vital role in the diagnosis and treatment of many neurological disorders. Recent years, deep learning technology has made great progress in the field of med…

cs.CV2022

A Weakly Supervised Learning Framework for Salient Object Detection via Hybrid Labels

Runmin Cong, Qi Qin, Chen Zhang +4

Fully-supervised salient object detection (SOD) methods have made great progress, but such methods often rely on a large number of pixel-level annotations, which are time-consuming…

cs.CV2018

Review of Visual Saliency Detection with Comprehensive Information

Runmin Cong, Jianjun Lei, Huazhu Fu +3

Visual saliency detection model simulates the human visual system to perceive the scene, and has been widely used in many vision tasks. With the acquisition technology development,…

cs.CV2017

Saliency Detection for Stereoscopic Images Based on Depth Confidence Analysis and Multiple Cues Fusion

Runmin Cong, Jianjun Lei, Changqing Zhang +3

Stereoscopic perception is an important part of human visual system that allows the brain to perceive depth. However, depth information has not been well explored in existing salie…

cs.RO2026

DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation

Zihao Xin, Wentong Li, Yixuan Jiang +4

Vision-and-Language Navigation (VLN) requires agents to follow long-horizon instructions and navigate complex 3D environments. However, existing approaches face two major challenge…

cs.CV2018

HSCS: Hierarchical Sparsity Based Co-saliency Detection for RGBD Images

Runmin Cong, Jianjun Lei, Huazhu Fu +3

Co-saliency detection aims to discover common and salient objects in an image group containing more than two relevant images. Moreover, depth information has been demonstrated to b…

eess.IV2025

Towards Robust and Generalizable Continuous Space-Time Video Super-Resolution with Events

Shuoyan Wei, Feng Li, Shengeng Tang +4

Continuous space-time video super-resolution (C-STVSR) has garnered increasing interest for its capability to reconstruct high-resolution and high-frame-rate videos at arbitrary sp…

cs.CV2025

The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation

Hao Fang, Runmin Cong, Xiankai Lu +2

Motion expression video segmentation is designed to segment objects in accordance with the input motion expressions. In contrast to the conventional Referring Video Object Segmenta…

cs.CV2024

Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object Detection

Runmin Cong, Hongyu Liu, Chen Zhang +4

By integrating complementary information from RGB image and depth map, the ability of salient object detection (SOD) for complex and challenging scenes can be improved. In recent y…

cs.CV2025

Semantic Concentration for Self-Supervised Dense Representations Learning

Peisong Wen, Qianqian Xu, Siran Dai +2

Recent advances in image-level self-supervised learning (SSL) have made significant progress, yet learning dense representations for patches remains challenging. Mainstream methods…

cs.CV2026

GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models

Bin Wang, Ruotong Hu, Wentong Li +5

Visual and textual soft prompt tuning can effectively improve the adaptability of Vision-Language Models (VLMs) in downstream tasks. However, fine-tuning on video tasks impairs the…

cs.CV2025

SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object Detection

Jia Lin, Xiaofei Zhou, Jiyuan Liu +4

Recently segment anything model (SAM) has attracted widespread concerns, and it is often treated as a vision foundation model for universal segmentation. Some researchers have atte…

cs.CV2025

PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild

Henghui Ding, Chang Liu, Nikhila Ravi +33

This report provides a comprehensive overview of the 4th Pixel-level Video Understanding in the Wild (PVUW) Challenge, held in conjunction with CVPR 2025. It summarizes the challen…

eess.IV2026

Unleashing Correlation and Continuity for Hyperspectral Reconstruction from RGB Images

Fuxiang Feng, Runmin Cong, Shoushui Wei +4

Reconstructing Hyperspectral Images (HSI) from RGB images can yield high spatial resolution HSI at a lower cost, demonstrating significant application potential. This paper reveals…

cs.CV2024

Size-invariance Matters: Rethinking Metrics and Losses for Imbalanced Multi-object Salient Object Detection

Feiran Li, Qianqian Xu, Shilong Bao +4

This paper explores the size-invariance of evaluation metrics in Salient Object Detection (SOD), especially when multiple targets of diverse sizes co-exist in the same image. We ob…

cs.CV2020

CoADNet: Collaborative Aggregation-and-Distribution Networks for Co-Salient Object Detection

Qijian Zhang, Runmin Cong, Junhui Hou +2

Co-Salient Object Detection (CoSOD) aims at discovering salient objects that repeatedly appear in a given query group containing two or more relevant images. One challenging issue…

cs.CV2025

Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model

Rui Xing, Runmin Cong, Yingying Wu +5

Understanding the dietary preferences of ancient societies and their evolution across periods and regions is crucial for revealing human-environment interactions. Seeds, as importa…

cs.AI2025

Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making

Liuxin Bao, Zhihao Peng, Xiaofei Zhou +3

Medical Decision-Making (MDM) is a complex process requiring substantial domain-specific expertise to effectively synthesize heterogeneous and complicated clinical information. Whi…

cs.CV2020

Global Context-Aware Progressive Aggregation Network for Salient Object Detection

Zuyao Chen, Qianqian Xu, Runmin Cong +1

Deep convolutional neural networks have achieved competitive performance in salient object detection, in which how to learn effective and comprehensive features plays a critical ro…

cs.CV2020

Dense Attention Fluid Network for Salient Object Detection in Optical Remote Sensing Images

Qijian Zhang, Runmin Cong, Chongyi Li +5

Despite the remarkable advances in visual saliency analysis for natural scene images (NSIs), salient object detection (SOD) for optical remote sensing images (RSIs) still remains a…

cs.CV2024

BlindDiff: Empowering Degradation Modelling in Diffusion Models for Blind Image Super-Resolution

Feng Li, Yixuan Wu, Zichao Liang +4

Diffusion models (DM) have achieved remarkable promise in image super-resolution (SR). However, most of them are tailored to solving non-blind inverse problems with fixed known deg…

cs.CV2025

Divide-and-Conquer Decoupled Network for Cross-Domain Few-Shot Segmentation

Runmin Cong, Anpeng Wang, Bin Wan +3

Cross-domain few-shot segmentation (CD-FSS) aims to tackle the dual challenge of recognizing novel classes and adapting to unseen domains with limited annotations. However, encoder…

eess.IV2025

Once-for-All: Controllable Generative Image Compression with Dynamic Granularity Adaptation

Anqi Li, Feng Li, Yuxi Liu +3

Although recent generative image compression methods have demonstrated impressive potential in optimizing the rate-distortion-perception trade-off, they still face the critical cha…

cs.CV2022

PSNet: Parallel Symmetric Network for Video Salient Object Detection

Runmin Cong, Weiyu Song, Jianjun Lei +3

For the video salient object detection (VSOD) task, how to excavate the information from the appearance modality and the motion modality has always been a topic of great concern. T…

cs.CV2026

Mastering Negation: Boosting Grounding Models via Grouped Opposition-Based Learning

Zesheng Yang, Xi Jiang, Bingzhang Hu +4

Current vision-language detection and grounding models predominantly focus on prompts with positive semantics and often struggle to accurately interpret and ground complex expressi…

cs.CV2017

Co-saliency Detection for RGBD Images Based on Multi-constraint Feature Matching and Cross Label Propagation

Runmin Cong, Jianjun Lei, Huazhu Fu +3

Co-saliency detection aims at extracting the common salient regions from an image group containing two or more relevant images. It is a newly emerging topic in computer vision comm…

cs.CV2017

An Iterative Co-Saliency Framework for RGBD Images

Runmin Cong, Jianjun Lei, Huazhu Fu +4

As a newly emerging and significant topic in computer vision community, co-saliency detection aims at discovering the common salient objects in multiple related images. The existin…

eess.IV2022

BCS-Net: Boundary, Context and Semantic for Automatic COVID-19 Lung Infection Segmentation from CT Images

Runmin Cong, Haowei Yang, Qiuping Jiang +5

The spread of COVID-19 has brought a huge disaster to the world, and the automatic segmentation of infection regions can help doctors to make diagnosis quickly and reduce workload.…

cs.CV2019

Nested Network with Two-Stream Pyramid for Salient Object Detection in Optical Remote Sensing Images

Chongyi Li, Runmin Cong, Junhui Hou +3

Arising from the various object types and scales, diverse imaging orientations, and cluttered backgrounds in optical remote sensing image (RSI), it is difficult to directly extend…

cs.CV2022

Learning Detail-Structure Alternative Optimization for Blind Super-Resolution

Feng Li, Yixuan Wu, Huihui Bai +3

Existing convolutional neural networks (CNN) based image super-resolution (SR) methods have achieved impressive performance on bicubic kernel, which is not valid to handle unknown…

cs.CV2026

G2HFNet: GeoGran-Aware Hierarchical Feature Fusion Network for Salient Object Detection in Optical Remote Sensing Images

Bin Wan, Runmin Cong, Xiaofei Zhou +3

Remote sensing images captured from aerial perspectives often exhibit significant scale variations and complex backgrounds, posing challenges for salient object detection (SOD). Ex…

cs.CV2020

RGB-D Salient Object Detection with Cross-Modality Modulation and Selection

Chongyi Li, Runmin Cong, Yongri Piao +2

We present an effective method to progressively integrate and refine the cross-modality complementarities for RGB-D salient object detection (SOD). The proposed network mainly solv…

cs.CV2026

Taming Real-World Space-Time Video Super-Resolution with One-Step Diffusion

Shuoyan Wei, Feng Li, Chen Zhou +3

Diffusion models have demonstrated exceptional success in video super-resolution (VSR), exhibiting powerful capabilities for generating fine-grained details. However, their potenti…

cs.CV2020

A Parallel Down-Up Fusion Network for Salient Object Detection in Optical Remote Sensing Images

Chongyi Li, Runmin Cong, Chunle Guo +4

The diverse spatial resolutions, various object types, scales and orientations, and cluttered backgrounds in optical remote sensing images (RSIs) challenge the current salient obje…

cs.CV2020

Learning Deep Interleaved Networks with Asymmetric Co-Attention for Image Restoration

Feng Li, Runmin Cong, Huihui Bai +3

Recently, convolutional neural network (CNN) has demonstrated significant success for image restoration (IR) tasks (e.g., image super-resolution, image deblurring, rain streak remo…

cs.CV2022

Global-and-Local Collaborative Learning for Co-Salient Object Detection

Runmin Cong, Ning Yang, Chongyi Li +4

The goal of co-salient object detection (CoSOD) is to discover salient objects that commonly appear in a query group containing two or more relevant images. Therefore, how to effec…

cs.CV2021

Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline

Lingzhi He, Hongguang Zhu, Feng Li +6

Depth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR)…

cs.CV2020

DPANet: Depth Potentiality-Aware Gated Attention Network for RGB-D Salient Object Detection

Zuyao Chen, Runmin Cong, Qianqian Xu +1

There are two main issues in RGB-D salient object detection: (1) how to effectively integrate the complementarity from the cross-modal RGB-D data; (2) how to prevent the contaminat…

cs.CL2026

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models

Zongji Yu, Wenshui Luo, Yiliu Sun +4

Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Group Relative Policy Optimizat…

cs.CV2025

NTIRE 2025 Challenge on Event-Based Image Deblurring: Methods and Results

Lei Sun, Andrea Alfarano, Peiqi Duan +85

This paper presents an overview of NTIRE 2025 the First Challenge on Event-Based Image Deblurring, detailing the proposed methodologies and corresponding results. The primary goal…

cs.CV2021

BridgeNet: A Joint Learning Network of Depth Map Super-Resolution and Monocular Depth Estimation

Qi Tang, Runmin Cong, Ronghui Sheng +4

Depth map super-resolution is a task with high practical application requirements in the industry. Existing color-guided depth map super-resolution methods usually necessitate an e…

cs.AI2026

TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning

Mingzu Liu, Hao Fang, Runmin Cong

Fine-Tuning-as-a-Service (FTaaS) facilitates the customization of Multimodal Large Language Models (MLLMs) but introduces critical backdoor risks via poisoned data. Existing defens…

cs.CV2025

Advancing Marine Research: UWSAM Framework and UIIS10K Dataset for Precise Underwater Instance Segmentation

Hua Li, Shijie Lian, Zhiyuan Li +5

With recent breakthroughs in large-scale modeling, the Segment Anything Model (SAM) has demonstrated significant potential in a variety of visual applications. However, due to the…

cs.CV2021

Cross-modality Discrepant Interaction Network for RGB-D Salient Object Detection

Chen Zhang, Runmin Cong, Qinwei Lin +4

The popularity and promotion of depth maps have brought new vigor and vitality into salient object detection (SOD), and a mass of RGB-D SOD algorithms have been proposed, mainly co…

cs.CV2026

M-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection

Jiyuan Liu, Jia Lin, Xiaofei Zhou +3

The Segment Anything Model 2 (SAM2) has emerged as a foundation model for universal segmentation. Owing to its generalizable visual representations, SAM2 has been successfully appl…

cs.LG2025

SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning

Jinpeng Chen, Runmin Cong, Yuzhi Zhao +4

Multimodal Continual Instruction Tuning (MCIT) aims to enable Multimodal Large Language Models (MLLMs) to incrementally learn new tasks without catastrophic forgetting. In this pap…

cs.CV2026

ICME 2026 Grand Challenge on Cross-Scenario Defect Detection and Fine-Grained Severity Grading for High-Precision Manufacturing

Wei Sun, Weixia Zhang, Linhan Cao +30

This paper presents the IEEE International Conference on Multimedia and Expo (ICME) 2026 Grand Challenge on Cross-Scenario Defect Detection and Fine-Grained Severity Grading for Hi…

cs.CV2025

Proto-FG3D: Prototype-based Interpretable Fine-Grained 3D Shape Classification

Shuxian Ma, Zihao Dong, Runmin Cong +2

Deep learning-based multi-view coarse-grained 3D shape classification has achieved remarkable success over the past decade, leveraging the powerful feature learning capabilities of…

cs.CV2025

From Sight to Insight: Unleashing Eye-Tracking in Weakly Supervised Video Salient Object Detection

Qi Qin, Runmin Cong, Gen Zhan +2

The eye-tracking video saliency prediction (VSP) task and video salient object detection (VSOD) task both focus on the most attractive objects in video and show the result in the f…

cs.CV2026

Beyond Global Scanning: Adaptive Visual State Space Modeling for Salient Object Detection in Optical Remote Sensing Images

Mengyu Ren, Yutong Li, Hua Li +2

Salient object detection (SOD) in optical remote sensing images (ORSIs) faces numerous challenges, including significant variations in target scales and low contrast between target…

cs.CV2024

UNINEXT-Cutie: The 1st Solution for LSVOS Challenge RVOS Track

Hao Fang, Feiyu Pan, Xiankai Lu +2

Referring video object segmentation (RVOS) relies on natural language expressions to segment target objects in video. In this year, LSVOS Challenge RVOS Track replaced the origin Y…

cs.CV2024

Strike a Balance in Continual Panoptic Segmentation

Jinpeng Chen, Runmin Cong, Yuxuan Luo +2

This study explores the emerging area of continual panoptic segmentation, highlighting three key balances. First, we introduce past-class backtrace distillation to balance the stab…

cs.CV2024

AIM 2024 Challenge on Video Saliency Prediction: Methods and Results

Andrey Moskalenko, Alexey Bryncev, Dmitry Vatolin +30

This paper reviews the Challenge on Video Saliency Prediction at AIM 2024. The goal of the participants was to develop a method for predicting accurate saliency maps for the provid…

eess.IV2025

Enhanced Quality Aware-Scalable Underwater Image Compression

Linwei Zhu, Junhao Zhu, Xu Zhang +4

Underwater imaging plays a pivotal role in marine exploration and ecological monitoring. However, it faces significant challenges of limited transmission bandwidth and severe disto…

cs.CV2022

CIR-Net: Cross-modality Interaction and Refinement for RGB-D Salient Object Detection

Runmin Cong, Qinwei Lin, Chen Zhang +4

Focusing on the issue of how to effectively capture and utilize cross-modality information in RGB-D salient object detection (SOD) task, we present a convolutional neural network (…

cs.CV2025

UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken

Runmin Cong, Zongji Yu, Hao Fang +2

Underwater Instance Segmentation (UIS) tasks are crucial for underwater complex scene detection. Mamba, as an emerging state space model with inherently linear complexity and globa…

eess.IV2022

Boundary Guided Semantic Learning for Real-time COVID-19 Lung Infection Segmentation System

Runmin Cong, Yumo Zhang, Ning Yang +6

The coronavirus disease 2019 (COVID-19) continues to have a negative impact on healthcare systems around the world, though the vaccines have been developed and national vaccination…

cs.CV2021

Underwater Image Enhancement via Medium Transmission-Guided Multi-Color Space Embedding

Chongyi Li, Saeed Anwar, Junhui Hou +3

Underwater images suffer from color casts and low contrast due to wavelength- and distance-dependent attenuation and scattering. To solve these two degradation issues, we present a…

cs.CV2024

Learning Hierarchical Color Guidance for Depth Map Super-Resolution

Runmin Cong, Ronghui Sheng, Hao Wu +5

Color information is the most commonly used prior knowledge for depth map super-resolution (DSR), which can provide high-frequency boundary guidance for detail restoration. However…

cs.CV2024

SDDNet: Style-guided Dual-layer Disentanglement Network for Shadow Detection

Runmin Cong, Yuchen Guan, Jinpeng Chen +3

Despite significant progress in shadow detection, current methods still struggle with the adverse impact of background color, which may lead to errors when shadows are present on c…

cs.CV2022

Multi-Projection Fusion and Refinement Network for Salient Object Detection in 360° Omnidirectional Image

Runmin Cong, Ke Huang, Jianjun Lei +3

Salient object detection (SOD) aims to determine the most visually attractive objects in an image. With the development of virtual reality technology, 360° omnidirectional image h…

eess.IV2023

Exploring Resolution Fields for Scalable Image Compression with Uncertainty Guidance

Dongyi Zhang, Feng Li, Man Liu +4

Recently, there are significant advancements in learning-based image compression methods surpassing traditional coding standards. Most of them prioritize achieving the best rate-di…

cs.CV2024

LSVOS Challenge Report: Large-scale Complex and Long Video Object Segmentation

Henghui Ding, Lingyi Hong, Chang Liu +30

Despite the promising performance of current video segmentation models on existing benchmarks, these models still struggle with complex scenes. In this paper, we introduce the 6th…

cs.CV2023

Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline

Tiantian Geng, Teng Wang, Jinming Duan +2

Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos o…

cs.CV2026

RSONet: Region-guided Selective Optimization Network for RGB-T Salient Object Detection

Bin Wan, Runmin Cong, Xiaofei Zhou +3

This paper focuses on the inconsistency in salient regions between RGB and thermal images. To address this issue, we propose the Region-guided Selective Optimization Network for RG…

cs.CV2026

Rethinking Conditional Generation for Underwater Salient Object Detection

Hua Li, Yongjie Weng, Yutong Li +3

Salient Object Detection in underwater images remains challenging due to low contrast, uneven illumination, and color distortion caused by scattering and absorption effects, which…

cs.CV2022

Does Thermal Really Always Matter for RGB-T Salient Object Detection?

Runmin Cong, Kepu Zhang, Chen Zhang +4

In recent years, RGB-T salient object detection (SOD) has attracted continuous attention, which makes it possible to identify salient objects in environments such as low light by i…

cs.CV2022

Stereo Superpixel Segmentation Via Decoupled Dynamic Spatial-Embedding Fusion Network

Hua Li, Junyan Liang, Ruiqi Wu +3

Stereo superpixel segmentation aims at grouping the discretizing pixels into perceptual regions through left and right views more collaboratively and efficiently. Existing superpix…

cs.CV2020

Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement

Chunle Guo, Chongyi Li, Jichang Guo +4

The paper presents a novel method, Zero-Reference Deep Curve Estimation (Zero-DCE), which formulates light enhancement as a task of image-specific curve estimation with a deep netw…

cs.CV2026

RDNet: Region Proportion-Aware Dynamic Adaptive Salient Object Detection Network in Optical Remote Sensing Images

Bin Wan, Runmin Cong, Xiaofei Zhou +3

Salient object detection (SOD) in remote sensing images faces significant challenges due to large variations in object sizes, the computational cost of self-attention mechanisms, a…

cs.CV2024

Frequency Perception Network for Camouflaged Object Detection

Runmin Cong, Mengyao Sun, Sanyi Zhang +3

Camouflaged object detection (COD) aims to accurately detect objects hidden in the surrounding environment. However, the existing COD methods mainly locate camouflaged objects in t…

eess.IV2020

NuI-Go: Recursive Non-Local Encoder-Decoder Network for Retinal Image Non-Uniform Illumination Removal

Chongyi Li, Huazhu Fu, Runmin Cong +2

Retinal images have been widely used by clinicians for early diagnosis of ocular diseases. However, the quality of retinal images is often clinically unsatisfactory due to eye lesi…

cs.CV2022

Bridging Component Learning with Degradation Modelling for Blind Image Super-Resolution

Yixuan Wu, Feng Li, Huihui Bai +3

Convolutional Neural Network (CNN)-based image super-resolution (SR) has exhibited impressive success on known degraded low-resolution (LR) images. However, this type of approach i…

cs.CV2023

You Can Mask More For Extremely Low-Bitrate Image Compression

Anqi Li, Feng Li, Jiaxin Han +6

Learned image compression (LIC) methods have experienced significant progress during recent years. However, these methods are primarily dedicated to optimizing the rate-distortion…

cs.CV2025

TDS-CLIP: Temporal Difference Side Network for Efficient VideoAction Recognition

Bin Wang, Wentong Li, Wenqian Wang +3

Recently, large-scale pre-trained vision-language models (e.g., CLIP), have garnered significant attention thanks to their powerful representative capabilities. This inspires resea…

cs.CV2022

RRNet: Relational Reasoning Network with Parallel Multi-scale Attention for Salient Object Detection in Optical Remote Sensing Images

Runmin Cong, Yumo Zhang, Leyuan Fang +3

Salient object detection (SOD) for optical remote sensing images (RSIs) aims at locating and extracting visually distinctive objects/regions from the optical RSIs. Despite some sal…

cs.CV2025

Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter

Zhiyang Chen, Chen Zhang, Hao Fang +1

Underwater instance segmentation (UIS), integrating pixel-level understanding and instance-level discrimination, is a pivotal technology in marine resource exploration and ecologic…

eess.IV2024

PUGAN: Physical Model-Guided Underwater Image Enhancement Using GAN with Dual-Discriminators

Runmin Cong, Wenyu Yang, Wei Zhang +4

Due to the light absorption and scattering induced by the water medium, underwater images usually suffer from some degradation problems, such as low contrast, color distortion, and…

cs.CV2024

Video Object Segmentation via SAM 2: The 4th Solution for LSVOS Challenge VOS Track

Feiyu Pan, Hao Fang, Runmin Cong +2

Video Object Segmentation (VOS) task aims to segmenting a particular object instance throughout the entire video sequence given only the object mask of the first frame. Recently, S…

cs.CV2024

Diving into Underwater: Segment Anything Model Guided Underwater Salient Instance Segmentation and A Large-scale Dataset

Shijie Lian, Ziyi Zhang, Hua Li +4

With the breakthrough of large models, Segment Anything Model (SAM) and its extensions have been attempted to apply in diverse tasks of computer vision. Underwater salient instance…