activity
20172021
most citedToward Transformer-Based Object Detection

141 citations · 145 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CV2021

Billion-Scale Pretraining with Vision Transformers for Multi-Task Visual Representations

Josh Beal, Hao-Yu Wu, Dong Huk Park +2

Large-scale pretraining of visual representations has led to state-of-the-art performance on a range of benchmark computer vision tasks, yet the benefits of these techniques at ext…

cs.CV2020141 cited

Toward Transformer-Based Object Detection

Josh Beal, Eric Kim, Eric Tzeng +3

Transformers have become the dominant model in natural language processing, owing to their ability to pretrain on massive amounts of data, then transfer to smaller, more specific t…

cs.CV2019

Learning a Unified Embedding for Visual Search at Pinterest

Andrew Zhai, Hao-Yu Wu, Eric Tzeng +2

At Pinterest, we utilize image embeddings throughout our search and recommendation systems to help our users navigate through visual content by powering experiences like browsing o…

cs.CV2019

Robust Change Captioning

Dong Huk Park, Trevor Darrell, Anna Rohrbach

Describing what has changed in a scene can be useful to a user, but only if generated text focuses on what is semantically relevant. It is thus important to distinguish distractors…

cs.AI2018

Multimodal Explanations: Justifying Decisions and Pointing to the Evidence

Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata +4

Deep models that are both effective and explainable are desirable in many settings; prior explainable models have been unimodal, offering either image-based visualization of attent…

cs.CV20174 cited

Attentive Explanations: Justifying Decisions and Pointing to the Evidence (Extended Abstract)

Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata +4

Deep models are the defacto standard in visual decision problems due to their impressive performance on a wide array of visual tasks. On the other hand, their opaqueness has led to…