papers

Publications (37)

cs.CV2025

Calibrating Undisciplined Over-Smoothing in Transformer for Weakly Supervised Semantic Segmentation

Lechao Cheng, Zerun Liu, Jingxuan He +3

Weakly supervised semantic segmentation (WSSS) has recently attracted considerable attention because it requires fewer annotations than fully supervised approaches, making it espec…

cs.CV2024

Revisiting the Power of Prompt for Visual Tuning

Yuzhu Wang, Lechao Cheng, Chaowei Fang +3

Visual prompt tuning (VPT) is a promising solution incorporating learnable prompt tokens to customize pre-trained models for downstream tasks. However, VPT and its variants often e…

cs.CV2022

Combating Noisy Labels in Long-Tailed Image Classification

Chaowei Fang, Lechao Cheng, Huiyan Qi +1

Most existing methods that cope with noisy labels usually assume that the class distributions are well balanced, which has insufficient capacity to deal with the practical scenario…

cs.CV2020

A Single Frame and Multi-Frame Joint Network for 360-degree Panorama Video Super-Resolution

Hongying Liu, Zhubo Ruan, Chaowei Fang +4

Spherical videos, also known as \ang{360} (panorama) videos, can be viewed with various virtual reality devices such as computers and head-mounted displays. They attract large amou…

cs.CV2025

Bridging Knowledge Gap Between Image Inpainting and Large-Area Visible Watermark Removal

Yicheng Leng, Chaowei Fang, Junye Chen +3

Visible watermark removal which involves watermark cleaning and background content restoration is pivotal to evaluate the resilience of watermarks. Existing deep neural network (DN…

cs.CV2024

Progressive Conservative Adaptation for Evolving Target Domains

Gangming Zhao, Chaoqi Chen, Wenhao He +5

Conventional domain adaptation typically transfers knowledge from a source domain to a stationary target domain. However, in many real-world cases, target data usually emerge seque…

cs.CV2022

Computer-aided Tuberculosis Diagnosis with Attribute Reasoning Assistance

Chengwei Pan, Gangming Zhao, Junjie Fang +6

Although deep learning algorithms have been intensively developed for computer-aided tuberculosis diagnosis (CTD), they mainly depend on carefully annotated datasets, leading to mu…

cs.CV2026

Diffusion Masked Pretraining for Dynamic Point Cloud

Zhuoyue Zhang, Jihua Zhu, Chaowei Fang +2

Dynamic point cloud pretraining is still dominated by masked reconstruction objectives. However, these objectives inherit two key limitations. Existing methods inject ground-truth…

cs.CV2021

MVCNet: Multiview Contrastive Network for Unsupervised Representation Learning for 3D CT Lesions

Penghua Zhai, Huaiwei Cong, Gangming Zhao +4

\emph{Objective and Impact Statement}. With the renaissance of deep learning, automatic diagnostic systems for computed tomography (CT) have achieved many successful applications.…

cs.MM2023

Removing Interference and Recovering Content Imaginatively for Visible Watermark Removal

Yicheng Leng, Chaowei Fang, Gen Li +2

Visible watermarks, while instrumental in protecting image copyrights, frequently distort the underlying content, complicating tasks like scene interpretation and image editing. Vi…

cs.CV2018

Piecewise Flat Embedding for Image Segmentation

Chaowei Fang, Zicheng Liao, Yizhou Yu

We introduce a new multi-dimensional nonlinear embedding -- Piecewise Flat Embedding (PFE) -- for image segmentation. Based on the theory of sparse signal recovery, piecewise flat…

cs.CV2022

Weakly Supervised Semantic Segmentation via Alternative Self-Dual Teaching

Dingwen Zhang, Wenyuan Zeng, Guangyu Guo +4

Current weakly supervised semantic segmentation (WSSS) frameworks usually contain the separated mask-refinement model and the main semantic region mining model. These approaches wo…

cs.CV2021

Deep Transformers for Fast Small Intestine Grounding in Capsule Endoscope Video

Xinkai Zhao, Chaowei Fang, Feng Gao +3

Capsule endoscopy is an evolutional technique for examining and diagnosing intractable gastrointestinal diseases. Because of the huge amount of data, analyzing capsule endoscope vi…

eess.IV2019

Globally Guided Progressive Fusion Network for 3D Pancreas Segmentation

Chaowei Fang, Guanbin Li, Chengwei Pan +2

Recently 3D volumetric organ segmentation attracts much research interest in medical image analysis due to its significance in computer aided diagnosis. This paper aims to address…

cs.CV2023

Variance-insensitive and Target-preserving Mask Refinement for Interactive Image Segmentation

Chaowei Fang, Ziyin Zhou, Junye Chen +3

Point-based interactive image segmentation can ease the burden of mask annotation in applications such as semantic segmentation and image editing. However, fully extracting the tar…

eess.IV2022

Deep 3D Vessel Segmentation based on Cross Transformer Network

Chengwei Pan, Baolian Qi, Gangming Zhao +4

The coronary microvascular disease poses a great threat to human health. Computer-aided analysis/diagnosis systems help physicians intervene in the disease at early stages, where 3…

eess.IV2020

Contralaterally Enhanced Networks for Thoracic Disease Detection

Gangming Zhao, Chaowei Fang, Guanbin Li +2

Identifying and locating diseases in chest X-rays are very challenging, due to the low visual contrast between normal and abnormal regions, and distortions caused by other overlapp…

cs.CV2022

Cross-Modality High-Frequency Transformer for MR Image Super-Resolution

Chaowei Fang, Dingwen Zhang, Liang Wang +3

Improving the resolution of magnetic resonance (MR) image data is critical to computer-aided diagnosis and brain function analysis. Higher resolution helps to capture more detailed…

cs.CV2020

Meta Corrupted Pixels Mining for Medical Image Segmentation

Jixin Wang, Sanping Zhou, Chaowei Fang +2

Deep neural networks have achieved satisfactory performance in piles of medical image analysis tasks. However the training of deep neural network requires a large amount of samples…

cs.CV2019

Self-Enhanced Convolutional Network for Facial Video Hallucination

Chaowei Fang, Guanbin Li, Xiaoguang Han +1

As a domain-specific super-resolution problem, facial image hallucination has enjoyed a series of breakthroughs thanks to the advances of deep convolutional neural networks. Howeve…

cs.CV2025

Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning

Fangwen Wu, Lechao Cheng, Shengeng Tang +4

Class-incremental learning (CIL) seeks to enable a model to sequentially learn new classes while retaining knowledge of previously learned ones. Balancing flexibility and stability…

cs.CV2025

Dual-domain Adaptation Networks for Realistic Image Super-resolution

Chaowei Fang, Bolin Fu, De Cheng +2

Realistic image super-resolution (SR) focuses on transforming real-world low-resolution (LR) images into high-resolution (HR) ones, handling more complex degradation patterns than…

cs.CV2023

Revisiting Long-tailed Image Classification: Survey and Benchmarks with New Evaluation Metrics

Chaowei Fang, Dingwen Zhang, Wen Zheng +4

Recently, long-tailed image classification harvests lots of research attention, since the data distribution is long-tailed in many real-world situations. Piles of algorithms are de…

eess.IV2022

Cross-level Contrastive Learning and Consistency Constraint for Semi-supervised Medical Image Segmentation

Xinkai Zhao, Chaowei Fang, De-Jun Fan +3

Semi-supervised learning (SSL), which aims at leveraging a few labeled images and a large number of unlabeled images for network training, is beneficial for relieving the burden of…

cs.CV2020

PNEN: Pyramid Non-Local Enhanced Networks

Feida Zhu, Chaowei Fang, Kai-Kuang Ma

Existing neural networks proposed for low-level image processing tasks are usually implemented by stacking convolution layers with limited kernel size. Every convolution layer mere…

cs.CV2020

Graph Neural Networks for UnsupervisedDomain Adaptation of Histopathological ImageAnalytics

Dou Xu, Chang Cai, Chaowei Fang +3

Annotating histopathological images is a time-consuming andlabor-intensive process, which requires broad-certificated pathologistscarefully examining large-scale whole-slide images…

cs.CV2026

Diffusion-based Layer-wise Semantic Reconstruction for Unsupervised Out-of-Distribution Detection

Ying Yang, De Cheng, Chaowei Fang +4

Unsupervised out-of-distribution (OOD) detection aims to identify out-of-domain data by learning only from unlabeled In-Distribution (ID) training samples, which is crucial for dev…

eess.IV2022

Incremental Cross-view Mutual Distillation for Self-supervised Medical CT Synthesis

Chaowei Fang, Liang Wang, Dingwen Zhang +3

Due to the constraints of the imaging device and high cost in operation time, computer tomography (CT) scans are usually acquired with low intra-slice resolution. Improving the int…

cs.CV2021

Densely Nested Top-Down Flows for Salient Object Detection

Chaowei Fang, Haibin Tian, Dingwen Zhang +3

With the goal of identifying pixel-wise salient object regions from each input image, salient object detection (SOD) has been receiving great attention in recent years. One kind of…

cs.CV2021

Trash to Treasure: Harvesting OOD Data with Cross-Modal Matching for Open-Set Semi-Supervised Learning

Junkai Huang, Chaowei Fang, Weikai Chen +5

Open-set semi-supervised learning (open-set SSL) investigates a challenging but practical scenario where out-of-distribution (OOD) samples are contained in the unlabeled data. Whil…

cs.CV2023

Progressive Feature Self-reinforcement for Weakly Supervised Semantic Segmentation

Jingxuan He, Lechao Cheng, Chaowei Fang +3

Compared to conventional semantic segmentation with pixel-level supervision, Weakly Supervised Semantic Segmentation (WSSS) with image-level labels poses the challenge that it alwa…

cs.CV2026

Tri-Efficient Transfer Learning for Point Cloud Videos

Yiding Sun, Dongxu Zhang, Jihua Zhu +6

While point cloud foundation models have significantly advanced point cloud video understanding, existing parameter-efficient fine-tuning (PEFT) methods still suffer from two criti…

cs.CV2023

Identity-Preserving Talking Face Generation with Landmark and Appearance Priors

Weizhi Zhong, Chaowei Fang, Yinqi Cai +4

Generating talking face videos from audio attracts lots of research interest. A few person-specific methods can generate vivid videos but require the target speaker's videos for tr…

cs.CV2022

Compound Batch Normalization for Long-tailed Image Classification

Lechao Cheng, Chaowei Fang, Dingwen Zhang +2

Significant progress has been made in learning image classification neural networks under long-tail data distribution using robust training algorithms such as data re-sampling, re-…

cs.MM2023

WMFormer++: Nested Transformer for Visible Watermark Removal via Implict Joint Learning

Dongjian Huo, Zehong Zhang, Hanjing Su +3

Watermarking serves as a widely adopted approach to safeguard media copyright. In parallel, the research focus has extended to watermark removal techniques, offering an adversarial…

cs.HC2026

Orchestrated Reality: From Role-Play to Living, Playable Game Worlds -- LLM-Driven World Simulation as a Parameterized-Action POMDP

Yuhang Huang, Chenmiao Li, Chaowei Fang

Many games rely on storytelling combined with systems that track levelling, NPC behaviour, and consequence simulation; bridging tightly-authored narrative with deeply-simulated wor…

cs.CV2026

PointCoT: A Multi-modal Benchmark for Explicit 3D Geometric Reasoning

Dongxu Zhang, Yiding Sun, Pengcheng Li +12

While Multimodal Large Language Models (MLLMs) demonstrate proficiency in 2D scenes, extending their perceptual intelligence to 3D point cloud understanding remains a significant c…