10 citations · 18 across the 4 of their papers we have counts for
7 papers
Enhanced Multimodal Representation Learning with Cross-modal KD
Mengxi Chen, Linyu Xing, Yu Wang +1
This paper explores the tasks of leveraging auxiliary modalities which are only available at training to enhance multimodal representation learning through cross-modal Knowledge Di…
Open-vocabulary Semantic Segmentation with Frozen Vision-Language Models
Chaofan Ma, Yuhuan Yang, Yanfeng Wang +2
When trained at a sufficient scale, self-supervised learning has exhibited a notable ability to solve a wide range of visual or language understanding tasks. In this paper, we inve…
Self-Supervised Masking for Unsupervised Anomaly Detection and Localization
Chaoqin Huang, Qinwei Xu, Yanfeng Wang +2
Recently, anomaly detection and localization in multimedia data have received significant attention among the machine learning community. In real-world applications such as medical…
Webpage Saliency Prediction with Two-stage Generative Adversarial Networks
Yu Li, Ya Zhang
Web page saliency prediction is a challenge problem in image transformation and computer vision. In this paper, we propose a new model combined with web page outline information to…
Multi-Scale Spatially-Asymmetric Recalibration for Image Classification
Yan Wang, Lingxi Xie, Siyuan Qiao +3
Convolution is spatially-symmetric, i.e., the visual features are independent of its position in the image, which limits its ability to utilize contextual cues for visual recogniti…
Chinese Typeface Transformation with Hierarchical Adversarial Network
Jie Chang, Yujun Gu, Ya Zhang
In this paper, we explore automated typeface generation through image style transfer which has shown great promise in natural image generation. Existing style transfer methods for…