most citedExpertPrompting: Instructing Large Language Models to be Distinguished Experts

56 citations · 186 across the 20 of their papers we have counts for

collaborators
Showing cs.CVShow all

15 papers · 1 filter

cs.CV2023

Dual-Stream Knowledge-Preserving Hashing for Unsupervised Video Retrieval

Pandeng Li, Hongtao Xie, Jiannan Ge +3

Unsupervised video hashing usually optimizes binary codes by learning to reconstruct input videos. Such reconstruction constraint spends much effort on frame-level temporal context…

cs.CV2023

Balanced Classification: A Unified Framework for Long-Tailed Object Detection

Tianhao Qi, Hongtao Xie, Pandeng Li +2

Conventional detectors suffer from performance degradation when dealing with long-tailed data due to a classification bias towards the majority head categories. In this paper, we c…

cs.CV2023★ 23 cited

MomentDiff: Generative Video Moment Retrieval from Random to Real

Pandeng Li, Chen-Wei Xie, Hongtao Xie +5

Video moment retrieval pursues an efficient and generalized solution to identify the specific temporal segments within an untrimmed video that correspond to a given language descri…

cs.CV2023★ 4 cited

DreamIdentity: Improved Editability for Efficient Face-identity Preserved Image Generation

Zhuowei Chen, Shancheng Fang, Wei Liu +4

While large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centric images, an intractable problem is how to preserve the face identity for con…

cs.CV2023★ 5 cited

Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization

Huan Ren, Wenfei Yang, Tianzhu Zhang +1

Weakly-supervised temporal action localization aims to localize and recognize actions in untrimmed videos with only video-level category labels during training. Without instance-le…

cs.CV2023★ 1 cited

Not All Image Regions Matter: Masked Vector Quantization for Autoregressive Image Generation

Mengqi Huang, Zhendong Mao, Quan Wang +1

Existing autoregressive models follow the two-stage generation paradigm that first learns a codebook in the latent space for image reconstruction and then completes the image gener…