3 citations · 6 across the 4 of their papers we have counts for
8 papers · 1 filter
Prior-enhanced Temporal Action Localization using Subject-aware Spatial Attention
Yifan Liu, Youbao Tang, Ning Zhang +2
Temporal action localization (TAL) aims to detect the boundary and identify the class of each action instance in a long untrimmed video. Current approaches treat video frames homog…
FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and Captioning
Suvir Mirchandani, Licheng Yu, Mengjiao Wang +4
Multimodal tasks in the fashion domain have significant potential for e-commerce, but involve challenging vision-and-language learning problems - e.g., retrieving a fashion item gi…
Unsupervised Vision-and-Language Pre-training via Retrieval-based Multi-Granular Alignment
Mingyang Zhou, Licheng Yu, Amanpreet Singh +3
Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require p…
Connecting What to Say With Where to Look by Modeling Human Attention Traces
Zihang Meng, Licheng Yu, Ning Zhang +4
We introduce a unified framework to jointly model images, text, and human attention traces. Our work is built on top of the recent Localized Narratives annotation framework [30], w…
Dynamic Kernel Distillation for Efficient Pose Estimation in Videos
Xuecheng Nie, Yuncheng Li, Linjie Luo +2
Existing video-based human pose estimation methods extensively apply large networks onto every frame in the video to localize body joints, which suffer high computational cost and…
Weakly Supervised Body Part Segmentation with Pose based Part Priors
Zhengyuan Yang, Yuncheng Li, Linjie Yang +2
Human body part segmentation refers to the task of predicting the semantic segmentation mask for each body part. Fully supervised body part segmentation methods achieve good perfor…