87 citations · 136 across the 4 of their papers we have counts for
6 papers · 1 filter
ContentCTR: Frame-level Live Streaming Click-Through Rate Prediction with Multimodal Transformer
Jiaxin Deng, Dong Shen, Shiyao Wang +4
In recent years, live streaming platforms have gained immense popularity as they allow users to broadcast their videos and interact in real-time with hosts and peers. Due to the dy…
Generation-Guided Multi-Level Unified Network for Video Grounding
Xing Cheng, Xiangyu Wu, Dong Shen +2
Video grounding aims to locate the timestamps best matching the query description within an untrimmed video. Prevalent methods can be divided into moment-level and clip-level frame…
MlTr: Multi-label Classification with Transformer
Xing Cheng, Hezheng Lin, Xiangyu Wu +5
The task of multi-label image classification is to recognize all the object labels presented in an image. Though advancing for years, small objects, similar objects and objects wit…
CAT: Cross Attention in Vision Transformer
Hezheng Lin, Xing Cheng, Xiangyu Wu +5
Since Transformer has found widespread use in NLP, the potential of Transformer in CV has been realized and has inspired many new approaches. However, the computation required for…
ES-Net: Erasing Salient Parts to Learn More in Re-Identification
Dong Shen, Shuai Zhao, Jinming Hu +3
As an instance-level recognition problem, re-identification (re-ID) requires models to capture diverse features. However, with continuous training, re-ID models pay more and more a…
Complementary Pseudo Labels For Unsupervised Domain Adaptation On Person Re-identification
Hao Feng, Minghao Chen, Jinming Hu +3
In recent years, supervised person re-identification (re-ID) models have received increasing studies. However, these models trained on the source domain always suffer dramatic perf…