activity
20172022
most citedOpen-vocabulary Semantic Segmentation with Frozen Vision-Language Models

10 citations · 18 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV2023

Enhanced Multimodal Representation Learning with Cross-modal KD

Mengxi Chen, Linyu Xing, Yu Wang +1

This paper explores the tasks of leveraging auxiliary modalities which are only available at training to enhance multimodal representation learning through cross-modal Knowledge Di…

cs.CV202210 cited

Open-vocabulary Semantic Segmentation with Frozen Vision-Language Models

Chaofan Ma, Yuhuan Yang, Yanfeng Wang +2

When trained at a sufficient scale, self-supervised learning has exhibited a notable ability to solve a wide range of visual or language understanding tasks. In this paper, we inve…

cs.CV2022

Self-Supervised Masking for Unsupervised Anomaly Detection and Localization

Chaoqin Huang, Qinwei Xu, Yanfeng Wang +2

Recently, anomaly detection and localization in multimedia data have received significant attention among the machine learning community. In real-world applications such as medical…

cs.CV2018

Webpage Saliency Prediction with Two-stage Generative Adversarial Networks

Yu Li, Ya Zhang

Web page saliency prediction is a challenge problem in image transformation and computer vision. In this paper, we propose a new model combined with web page outline information to…

cs.CV2018

Multi-Scale Spatially-Asymmetric Recalibration for Image Classification

Yan Wang, Lingxi Xie, Siyuan Qiao +3

Convolution is spatially-symmetric, i.e., the visual features are independent of its position in the image, which limits its ability to utilize contextual cues for visual recogniti…

cs.CV20177 cited

Chinese Typeface Transformation with Hierarchical Adversarial Network

Jie Chang, Yujun Gu, Ya Zhang

In this paper, we explore automated typeface generation through image style transfer which has shown great promise in natural image generation. Existing style transfer methods for…