71 citations · 91 across the 13 of their papers we have counts for
5 papers · 1 filter
Textual Inversion and Self-supervised Refinement for Radiology Report Generation
Yuanjiang Luo, Hongxiang Li, Xuan Wu +6
Existing mainstream approaches follow the encoder-decoder paradigm for generating radiology reports. They focus on improving the network structure of encoders and decoders, which l…
Video Referring Expression Comprehension via Transformer with Content-conditioned Query
Ji Jiang, Meng Cao, Tengtao Song +3
Video Referring Expression Comprehension (REC) aims to localize a target object in videos based on the queried natural language. Recent improvements in video REC have been made usi…
RGI: robust GAN-inversion for mask-free image inpainting and unsupervised pixel-wise anomaly detection
Shancong Mou, Xiaoyi Gu, Meng Cao +4
Generative adversarial networks (GANs), trained on a large-scale image dataset, can be a good approximator of the natural image manifold. GAN-inversion, using a pre-trained generat…
Correspondence Matters for Video Referring Expression Comprehension
Meng Cao, Ji Jiang, Long Chen +1
We investigate the problem of video Referring Expression Comprehension (REC), which aims to localize the referent objects described in the sentence to visual regions in the video f…
LocVTP: Video-Text Pre-training for Temporal Localization
Meng Cao, Tianyu Yang, Junwu Weng +3
Video-Text Pre-training (VTP) aims to learn transferable representations for various downstream tasks from large-scale web videos. To date, almost all existing VTP methods are limi…