6 citations · 13 across the 4 of their papers we have counts for
6 papers
CPL: Counterfactual Prompt Learning for Vision and Language Models
Xuehai He, Diji Yang, Weixi Feng +7
Prompt tuning is a new few-shot transfer learning technique that only tunes the learnable prompt for pre-trained vision and language models such as CLIP. However, existing prompt t…
Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations
Wanrong Zhu, Xin Eric Wang, Pradyumna Narayana +3
A major challenge in visually grounded language generation is to build robust benchmark datasets and models that can generalize well in real-world settings. To do this, it is criti…
Leveraging Organizational Resources to Adapt Models to New Data Modalities
Sahaana Suri, Raghuveer Chanda, Neslihan Bulut +7
As applications in large organizations evolve, the machine learning (ML) models that power them must adapt the same predictive tasks to newly arising data modalities (e.g., a new v…
Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation
Wanrong Zhu, Xin Eric Wang, Tsu-Jui Fu +5
One of the most challenging topics in Natural Language Processing (NLP) is visually-grounded language understanding and reasoning. Outdoor vision-and-language navigation (VLN) is s…
Multi-Image Summarization: Textual Summary from a Set of Cohesive Images
Nicholas Trieu, Sebastian Goodman, Pradyumna Narayana +2
Multi-sentence summarization is a well studied problem in NLP, while generating image descriptions for a single image is a well studied problem in Computer Vision. However, for app…
HUSE: Hierarchical Universal Semantic Embeddings
Pradyumna Narayana, Aniket Pednekar, Abishek Krishnamoorthy +2
There is a recent surge of interest in cross-modal representation learning corresponding to images and text. The main challenge lies in mapping images and text to a shared latent s…