84 citations · 153 across the 7 of their papers we have counts for
12 papers · 1 filter
X-HRNet: Towards Lightweight Human Pose Estimation with Spatially Unidimensional Self-Attention
Yixuan Zhou, Xuanhan Wang, Xing Xu +2
High-resolution representation is necessary for human pose estimation to achieve high performance, and the ensuing problem is high computational complexity. In particular, predomin…
MSFlow: Multi-Scale Flow-based Framework for Unsupervised Anomaly Detection
Yixuan Zhou, Xing Xu, Jingkuan Song +2
Unsupervised anomaly detection (UAD) attracts a lot of research interest and drives widespread applications, where only anomaly-free samples are available for training. Some UAD ap…
Unifying Two-Stream Encoders with Transformers for Cross-Modal Retrieval
Yi Bin, Haoxuan Li, Yahui Xu +3
Most existing cross-modal retrieval methods employ two-stream encoders with different architectures for images and texts, \textit{e.g.}, CNN for images and RNN/Transformer for text…
Faster Video Moment Retrieval with Point-Level Supervision
Xun Jiang, Zailei Zhou, Xing Xu +3
Video Moment Retrieval (VMR) aims at retrieving the most relevant events from an untrimmed video with natural language queries. Existing VMR methods suffer from two defects: (1) ma…
From General to Specific: Informative Scene Graph Generation via Balance Adjustment
Yuyu Guo, Lianli Gao, Xuanhan Wang +5
The scene graph generation (SGG) task aims to detect visual relationship triplets, i.e., subject, predicate, object, in an image, providing a structural vision layout for scene und…
Feature Space Targeted Attacks by Statistic Alignment
Lianli Gao, Yaya Cheng, Qilong Zhang +2
By adding human-imperceptible perturbations to images, DNNs can be easily fooled. As one of the mainstream methods, feature space targeted attacks perturb images by modulating thei…