56 citations · 186 across the 20 of their papers we have counts for
15 papers · 1 filter
Dual-Stream Knowledge-Preserving Hashing for Unsupervised Video Retrieval
Pandeng Li, Hongtao Xie, Jiannan Ge +3
Unsupervised video hashing usually optimizes binary codes by learning to reconstruct input videos. Such reconstruction constraint spends much effort on frame-level temporal context…
Balanced Classification: A Unified Framework for Long-Tailed Object Detection
Tianhao Qi, Hongtao Xie, Pandeng Li +2
Conventional detectors suffer from performance degradation when dealing with long-tailed data due to a classification bias towards the majority head categories. In this paper, we c…
MomentDiff: Generative Video Moment Retrieval from Random to Real
Pandeng Li, Chen-Wei Xie, Hongtao Xie +5
Video moment retrieval pursues an efficient and generalized solution to identify the specific temporal segments within an untrimmed video that correspond to a given language descri…
DreamIdentity: Improved Editability for Efficient Face-identity Preserved Image Generation
Zhuowei Chen, Shancheng Fang, Wei Liu +4
While large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centric images, an intractable problem is how to preserve the face identity for con…
Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization
Huan Ren, Wenfei Yang, Tianzhu Zhang +1
Weakly-supervised temporal action localization aims to localize and recognize actions in untrimmed videos with only video-level category labels during training. Without instance-le…
Not All Image Regions Matter: Masked Vector Quantization for Autoregressive Image Generation
Mengqi Huang, Zhendong Mao, Quan Wang +1
Existing autoregressive models follow the two-stage generation paradigm that first learns a codebook in the latent space for image reconstruction and then completes the image gener…