173 citations · 175 across the 4 of their papers we have counts for
4 papers
Likelihood-Based Text-to-Image Evaluation with Patch-Level Perceptual and Semantic Credit Assignment
Qi Chen, Chaorui Deng, Zixiong Huang +3
Text-to-image synthesis has made encouraging progress and attracted lots of public attention recently. However, popular evaluation metrics in this area, like the Inception Score an…
Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval
Chaorui Deng, Qi Chen, Pengda Qin +2
In text-video retrieval, recent works have benefited from the powerful learning capabilities of pre-trained text-image foundation models (e.g., CLIP) by adapting them to the video…
Learning to Dub Movies via Hierarchical Prosody Models
Gaoxiang Cong, Liang Li, Yuankai Qi +6
Given a piece of text, a video clip and a reference audio, the movie dubbing (also known as visual voice clone V2C) task aims to generate speeches that match the speaker's emotion…
Medical Visual Question Answering: A Survey
Zhihong Lin, Donghao Zhang, Qingyi Tao +5
Medical Visual Question Answering~(VQA) is a combination of medical artificial intelligence and popular VQA challenges. Given a medical image and a clinically relevant question in…