activity
20212024
most citedLearning Generative Vision Transformer with Energy-Based Latent Space for Saliency Prediction

45 citations · 53 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

Learning Gaussian Representation for Eye Fixation Prediction

Peipei Song, Jing Zhang, Piotr Koniusz +1

Existing eye fixation prediction methods perform the mapping from input images to the corresponding dense fixation maps generated from raw fixation points. However, due to the stoc…

cs.CV2023

Decompose Semantic Shifts for Composed Image Retrieval

Xingyu Yang, Daqing Liu, Heng Zhang +3

Composed image retrieval is a type of image retrieval task where the user provides a reference image as a starting point and specifies a text on how to shift from the starting poin…

cs.CV20235 cited

ESSAformer: Efficient Transformer for Hyperspectral Image Super-resolution

Mingjin Zhang, Chi Zhang, Qiming Zhang +3

Single hyperspectral image super-resolution (single-HSI-SR) aims to restore a high-resolution hyperspectral image from a low-resolution observation. However, the prevailing CNN-bas…

cs.CV20231 cited

Human-imperceptible, Machine-recognizable Images

Fusheng Hao, Fengxiang He, Yikai Wang +4

Massive human-related data is collected to train neural networks for computer vision tasks. A major conflict is exposed relating to software engineers between better developing AI…

cs.CV20232 cited

Mutual Information Regularization for Weakly-supervised RGB-D Salient Object Detection

Aixuan Li, Yuxin Mao, Jing Zhang +1

In this paper, we present a weakly-supervised RGB-D salient object detection model via scribble supervision. Specifically, as a multimodal learning task, we focus on effective mult…

cs.CV2021

Siamese Network with Interactive Transformer for Video Object Segmentation

Meng Lan, Jing Zhang, Fengxiang He +1

Semi-supervised video object segmentation (VOS) refers to segmenting the target object in remaining frames given its annotation in the first frame, which has been actively studied…