most citedUnifying Two-Stream Encoders with Transformers for Cross-Modal Retrieval

33 citations · 63 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2023

X-HRNet: Towards Lightweight Human Pose Estimation with Spatially Unidimensional Self-Attention

Yixuan Zhou, Xuanhan Wang, Xing Xu +2

High-resolution representation is necessary for human pose estimation to achieve high performance, and the ensuing problem is high computational complexity. In particular, predomin…

cs.CV20233 cited

MSFlow: Multi-Scale Flow-based Framework for Unsupervised Anomaly Detection

Yixuan Zhou, Xing Xu, Jingkuan Song +2

Unsupervised anomaly detection (UAD) attracts a lot of research interest and drives widespread applications, where only anomaly-free samples are available for training. Some UAD ap…

cs.CV20231 cited

ImbSAM: A Closer Look at Sharpness-Aware Minimization in Class-Imbalanced Recognition

Yixuan Zhou, Yi Qu, Xing Xu +1

Class imbalance is a common challenge in real-world recognition tasks, where the majority of classes have few samples, also known as tail classes. We address this challenge with th…

cs.CV202333 cited

Unifying Two-Stream Encoders with Transformers for Cross-Modal Retrieval

Yi Bin, Haoxuan Li, Yahui Xu +3

Most existing cross-modal retrieval methods employ two-stream encoders with different architectures for images and texts, \textit{e.g.}, CNN for images and RNN/Transformer for text…

cs.CV2023

Do-GOOD: Towards Distribution Shift Evaluation for Pre-Trained Visual Document Understanding Models

Jiabang He, Yi Hu, Lei Wang +4

Numerous pre-training techniques for visual document understanding (VDU) have recently shown substantial improvements in performance across a wide range of document tasks. However,…

cs.CV20231 cited

Faster Video Moment Retrieval with Point-Level Supervision

Xun Jiang, Zailei Zhou, Xing Xu +3

Video Moment Retrieval (VMR) aims at retrieving the most relevant events from an untrimmed video with natural language queries. Existing VMR methods suffer from two defects: (1) ma…