activity
20192022
most citedLeveraging Organizational Resources to Adapt Models to New Data Modalities

6 citations · 13 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CV20223 cited

CPL: Counterfactual Prompt Learning for Vision and Language Models

Xuehai He, Diji Yang, Weixi Feng +7

Prompt tuning is a new few-shot transfer learning technique that only tunes the learnable prompt for pre-trained vision and language models such as CLIP. However, existing prompt t…

cs.CL2020

Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations

Wanrong Zhu, Xin Eric Wang, Pradyumna Narayana +3

A major challenge in visually grounded language generation is to build robust benchmark datasets and models that can generalize well in real-world settings. To do this, it is criti…

cs.LG20206 cited

Leveraging Organizational Resources to Adapt Models to New Data Modalities

Sahaana Suri, Raghuveer Chanda, Neslihan Bulut +7

As applications in large organizations evolve, the machine learning (ML) models that power them must adapt the same predictive tasks to newly arising data modalities (e.g., a new v…

cs.CL2020

Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation

Wanrong Zhu, Xin Eric Wang, Tsu-Jui Fu +5

One of the most challenging topics in Natural Language Processing (NLP) is visually-grounded language understanding and reasoning. Outdoor vision-and-language navigation (VLN) is s…

cs.CV20201 cited

Multi-Image Summarization: Textual Summary from a Set of Cohesive Images

Nicholas Trieu, Sebastian Goodman, Pradyumna Narayana +2

Multi-sentence summarization is a well studied problem in NLP, while generating image descriptions for a single image is a well studied problem in Computer Vision. However, for app…

cs.CV20193 cited

HUSE: Hierarchical Universal Semantic Embeddings

Pradyumna Narayana, Aniket Pednekar, Abishek Krishnamoorthy +2

There is a recent surge of interest in cross-modal representation learning corresponding to images and text. The main challenge lies in mapping images and text to a shared latent s…