2 citations · 4 across the 3 of their papers we have counts for
4 papers · 1 filter
CLIP also Understands Text: Prompting CLIP for Phrase Understanding
An Yan, Jiacheng Li, Wanrong Zhu +3
Contrastive Language-Image Pretraining (CLIP) efficiently learns visual concepts by pre-training with natural language supervision. CLIP and its visual encoder have been explored o…
Weakly Supervised Contrastive Learning for Chest X-Ray Report Generation
An Yan, Zexue He, Xing Lu +5
Radiology report generation aims at generating descriptive text from radiology images automatically, which may present an opportunity to improve radiology reporting and interpretat…
Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation
Wanrong Zhu, Xin Eric Wang, Tsu-Jui Fu +5
One of the most challenging topics in Natural Language Processing (NLP) is visually-grounded language understanding and reasoning. Outdoor vision-and-language navigation (VLN) is s…
Cross-Lingual Vision-Language Navigation
An Yan, Xin Eric Wang, Jiangtao Feng +2
Commanding a robot to navigate with natural language instructions is a long-term goal for grounded language understanding and robotics. But the dominant language is English, accord…