141 citations · 145 across the 3 of their papers we have counts for
6 papers
Billion-Scale Pretraining with Vision Transformers for Multi-Task Visual Representations
Josh Beal, Hao-Yu Wu, Dong Huk Park +2
Large-scale pretraining of visual representations has led to state-of-the-art performance on a range of benchmark computer vision tasks, yet the benefits of these techniques at ext…
Toward Transformer-Based Object Detection
Josh Beal, Eric Kim, Eric Tzeng +3
Transformers have become the dominant model in natural language processing, owing to their ability to pretrain on massive amounts of data, then transfer to smaller, more specific t…
Learning a Unified Embedding for Visual Search at Pinterest
Andrew Zhai, Hao-Yu Wu, Eric Tzeng +2
At Pinterest, we utilize image embeddings throughout our search and recommendation systems to help our users navigate through visual content by powering experiences like browsing o…
Robust Change Captioning
Dong Huk Park, Trevor Darrell, Anna Rohrbach
Describing what has changed in a scene can be useful to a user, but only if generated text focuses on what is semantically relevant. It is thus important to distinguish distractors…
Multimodal Explanations: Justifying Decisions and Pointing to the Evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata +4
Deep models that are both effective and explainable are desirable in many settings; prior explainable models have been unimodal, offering either image-based visualization of attent…
Attentive Explanations: Justifying Decisions and Pointing to the Evidence (Extended Abstract)
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata +4
Deep models are the defacto standard in visual decision problems due to their impressive performance on a wide array of visual tasks. On the other hand, their opaqueness has led to…