67 citations · 135 across the 12 of their papers we have counts for
14 papers · 1 filter
March in Chat: Interactive Prompting for Remote Embodied Referring Expression
Yanyuan Qiao, Yuankai Qi, Zheng Yu +2
Many Vision-and-Language Navigation (VLN) tasks have been proposed in recent years, from room-based to object-based and indoor to outdoor. The REVERIE (Remote Embodied Referring Ex…
AerialVLN: Vision-and-Language Navigation for UAVs
Shubo Liu, Hongsheng Zhang, Yuankai Qi +3
Recently emerged Vision-and-Language Navigation (VLN) tasks have drawn significant attention in both computer vision and natural language processing communities. Existing VLN tasks…
Teacher Agent: A Knowledge Distillation-Free Framework for Rehearsal-based Video Incremental Learning
Shengqin Jiang, Yaoyu Fang, Haokui Zhang +4
Rehearsal-based video incremental learning often employs knowledge distillation to mitigate catastrophic forgetting of previously learned data. However, this method faces two major…
Progressive Multi-resolution Loss for Crowd Counting
Ziheng Yan, Yuankai Qi, Guorong Li +4
Crowd counting is usually handled in a density map regression fashion, which is supervised via a L2 loss between the predicted density map and ground truth. To effectively regulate…
Exploiting Completeness and Uncertainty of Pseudo Labels for Weakly Supervised Video Anomaly Detection
Chen Zhang, Guorong Li, Yuankai Qi +4
Weakly supervised video anomaly detection aims to identify abnormal events in videos using only video-level labels. Recently, two-stage self-training methods have achieved signific…
Consistency-Aware Anchor Pyramid Network for Crowd Localization
Xinyan Liu, Guorong Li, Yuankai Qi +4
Crowd localization aims to predict the spatial position of humans in a crowd scenario. We observe that the performance of existing methods is challenged from two aspects: (i) ranki…