4 citations · 11 across the 5 of their papers we have counts for
4 papers · 1 filter
Hypotheses Tree Building for One-Shot Temporal Sentence Localization
Daizong Liu, Xiang Fang, Pan Zhou +3
Given an untrimmed video, temporal sentence localization (TSL) aims to localize a specific segment according to a given sentence query. Though respectable works have made decent ac…
Frido: Feature Pyramid Diffusion for Complex Scene Image Synthesis
Wan-Cyuan Fan, Yen-Chun Chen, Dongdong Chen +3
Diffusion models (DMs) have shown great potential for high-quality image synthesis. However, when it comes to producing images with complex scenes, how to properly describe both im…
Bottom-Up 2D Pose Estimation via Dual Anatomical Centers for Small-Scale Persons
Yu Cheng, Yihao Ai, Bo Wang +2
In multi-person 2D pose estimation, the bottom-up methods simultaneously predict poses for all persons, and unlike the top-down methods, do not rely on human detection. However, th…
Learning Visual Representation from Modality-Shared Contrastive Language-Image Pre-training
Haoxuan You, Luowei Zhou, Bin Xiao +5
Large-scale multi-modal contrastive pre-training has demonstrated great utility to learn transferable features for a range of downstream tasks by mapping multiple modalities into a…