activity
20202022
most citedFine-grained Semantic Alignment Network for Weakly Supervised Temporal Language Grounding

19 citations · 40 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2023

Text-Only Training for Visual Storytelling

Yuechen Wang, Wengang Zhou, Zhenbo Lu +1

Visual storytelling aims to generate a narrative based on a sequence of images, necessitating both vision-language alignment and coherent story generation. Most existing solutions…

cs.CV202219 cited

Fine-grained Semantic Alignment Network for Weakly Supervised Temporal Language Grounding

Yuechen Wang, Wengang Zhou, Houqiang Li

Temporal language grounding (TLG) aims to localize a video segment in an untrimmed video based on a natural language description. To alleviate the expensive cost of manual annotati…

cs.CV20221 cited

Geometric Representation Learning for Document Image Rectification

Hao Feng, Wengang Zhou, Jiajun Deng +2

In document image rectification, there exist rich geometric constraints between the distorted image and the ground truth one. However, such geometric constraints are largely ignore…

cs.CV20212 cited

SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition

Hezhen Hu, Weichao Zhao, Wengang Zhou +2

Hand gesture serves as a critical role in sign language. Current deep-learning-based sign language recognition (SLR) methods may suffer insufficient interpretability and overfittin…

cs.CV202111 cited

Weakly Supervised Temporal Adjacent Network for Language Grounding

Yuechen Wang, Jiajun Deng, Wengang Zhou +1

Temporal language grounding (TLG) is a fundamental and challenging problem for vision and language understanding. Existing methods mainly focus on fully supervised setting with tem…