activity
20212023
most citedLAVT: Language-Aware Vision Transformer for Referring Image Segmentation

22 citations · 47 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV20232 cited

Extending Multi-modal Contrastive Representations

Zehan Wang, Ziang Zhang, Luping Liu +4

Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high d…

cs.CV20231 cited

CLIP3Dstyler: Language Guided 3D Arbitrary Neural Style Transfer

Ming Gao, YanWu Xu, Yang Zhao +3

In this paper, we propose a novel language-guided 3D arbitrary neural style transfer method (CLIP3Dstyler). We aim at stylizing any 3D scene with an arbitrary style from a text des…

cs.CL2023

Multi-Teacher Knowledge Distillation For Text Image Machine Translation

Cong Ma, Yaping Zhang, Mei Tu +3

Text image machine translation (TIMT) has been widely used in various real-world applications, which translates source language texts in images into another target language sentenc…

cs.CL2023

E2TIMT: Efficient and Effective Modal Adapter for Text Image Machine Translation

Cong Ma, Yaping Zhang, Mei Tu +3

Text image machine translation (TIMT) aims to translate texts embedded in images from one source language to another target language. Existing methods, both two-stage cascade and o…

cs.CV20231 cited

DATE: Domain Adaptive Product Seeker for E-commerce

Haoyuan Li, Hao Jiang, Tao Jin +5

Product Retrieval (PR) and Grounding (PG), aiming to seek image and object-level products respectively according to a textual query, have attracted great interest recently for bett…

cs.CV202321 cited

Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models

Xuhui Jia, Yang Zhao, Kelvin C. K. Chan +6

This paper proposes a method for generating images of customized objects specified by users. The method is based on a general framework that bypasses the lengthy optimization requi…