most citedLAVT: Language-Aware Vision Transformer for Referring Image Segmentation

22 citations · 23 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2023

Deep Fusion Transformer Network with Weighted Vector-Wise Keypoints Voting for Robust 6D Object Pose Estimation

Jun Zhou, Kai Chen, Linlin Xu +2

One critical challenge in 6D object pose estimation from a single RGBD image is efficient integration of two different modalities, i.e., color and depth. In this work, we tackle th…

cs.CV2023

Improving Handwritten OCR with Training Samples Generated by Glyph Conditional Denoising Diffusion Probabilistic Model

Haisong Ding, Bozhi Luan, Dongnan Gui +2

Constructing a highly accurate handwritten OCR system requires large amounts of representative training data, which is both time-consuming and expensive to collect. To mitigate the…

cs.CV2023

Zero-shot Generation of Training Data with Denoising Diffusion Probabilistic Model for Handwritten Chinese Character Recognition

Dongnan Gui, Kai Chen, Haisong Ding +1

There are more than 80,000 character categories in Chinese while most of them are rarely used. To build a high performance handwritten Chinese character recognition (HCCR) system s…

cs.CV20231 cited

Semantics-Aware Dynamic Localization and Refinement for Referring Image Segmentation

Zhao Yang, Jiaqi Wang, Yansong Tang +3

Referring image segmentation segments an image from a language expression. With the aim of producing high-quality masks, existing methods often adopt iterative learning approaches…

cs.CV202122 cited

LAVT: Language-Aware Vision Transformer for Referring Image Segmentation

Zhao Yang, Jiaqi Wang, Yansong Tang +3

Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image. One of the key cha…