63 citations · 79 across the 15 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
Yu Zhao, Hao Fei, Xiangtai Li +6
In the visual spatial understanding (VSU) area, spatial image-to-text (SI2T) and spatial text-to-image (ST2I) are two fundamental tasks that appear in dual form. Existing methods f…
cs.CV2024★ 63 cited
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
Hao Fei, Shengqiong Wu, Meishan Zhang +3
While pre-training large-scale video-language models (VLMs) has shown remarkable potential for various downstream video-language tasks, existing VLMs can still suffer from certain…