413 citations · 1.3k across the 37 of their papers we have counts for
26 papers · 1 filter
Foreground-Action Consistency Network for Weakly Supervised Temporal Action Localization
Linjiang Huang, Liang Wang, Hongsheng Li
As a challenging task of high-level video understanding, weakly supervised temporal action localization has been attracting increasing attention. With only video annotations, most…
Neighbor-view Enhanced Model for Vision and Language Navigation
Dong An, Yuankai Qi, Yan Huang +3
Vision and Language Navigation (VLN) requires an agent to navigate to a target location by following natural language instructions. Most of existing works represent a navigation ca…
Adaptive Dilated Convolution For Human Pose Estimation
Zhengxiong Luo, Zhicheng Wang, Yan Huang +3
Most existing human pose estimation (HPE) methods exploit multi-scale information by fusing feature maps of four different spatial sizes, \ie , , , and of th…
CMF: Cascaded Multi-model Fusion for Referring Image Segmentation
Jianhua Yang, Yan Huang, Zhanyu Ma +1
In this work, we address the task of referring image segmentation (RIS), which aims at predicting a segmentation mask for the object described by a natural language expression. Mos…
Few-Shot Learning with Part Discovery and Augmentation from Unlabeled Images
Wentao Chen, Chenyang Si, Wei Wang +3
Few-shot learning is a challenging task since only few instances are given for recognizing an unseen class. One way to alleviate this problem is to acquire a strong inductive bias…
End-to-end Alternating Optimization for Blind Super Resolution
Zhengxiong Luo, Yan Huang, Shang Li +2
Previous methods decompose the blind super-resolution (SR) problem into two sequential steps: \textit{i}) estimating the blur kernel from given low-resolution (LR) image and \texti…