22 citations · 47 across the 8 of their papers we have counts for
8 papers
Extending Multi-modal Contrastive Representations
Zehan Wang, Ziang Zhang, Luping Liu +4
Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high d…
CLIP3Dstyler: Language Guided 3D Arbitrary Neural Style Transfer
Ming Gao, YanWu Xu, Yang Zhao +3
In this paper, we propose a novel language-guided 3D arbitrary neural style transfer method (CLIP3Dstyler). We aim at stylizing any 3D scene with an arbitrary style from a text des…
Multi-Teacher Knowledge Distillation For Text Image Machine Translation
Cong Ma, Yaping Zhang, Mei Tu +3
Text image machine translation (TIMT) has been widely used in various real-world applications, which translates source language texts in images into another target language sentenc…
E2TIMT: Efficient and Effective Modal Adapter for Text Image Machine Translation
Cong Ma, Yaping Zhang, Mei Tu +3
Text image machine translation (TIMT) aims to translate texts embedded in images from one source language to another target language. Existing methods, both two-stage cascade and o…
DATE: Domain Adaptive Product Seeker for E-commerce
Haoyuan Li, Hao Jiang, Tao Jin +5
Product Retrieval (PR) and Grounding (PG), aiming to seek image and object-level products respectively according to a textual query, have attracted great interest recently for bett…
Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models
Xuhui Jia, Yang Zhao, Kelvin C. K. Chan +6
This paper proposes a method for generating images of customized objects specified by users. The method is based on a general framework that bypasses the lengthy optimization requi…