141 citations · 153 across the 4 of their papers we have counts for
4 papers · 1 filter
PinCLIP: Large-scale Foundational Multimodal Representation at Pinterest
Josh Beal, Eric Kim, Jinfeng Rao +3
While multi-modal Visual Language Models (VLMs) have demonstrated significant success across various domains, the integration of VLMs into recommendation and retrieval systems rema…
Billion-Scale Pretraining with Vision Transformers for Multi-Task Visual Representations
Josh Beal, Hao-Yu Wu, Dong Huk Park +2
Large-scale pretraining of visual representations has led to state-of-the-art performance on a range of benchmark computer vision tasks, yet the benefits of these techniques at ext…
Toward Transformer-Based Object Detection
Josh Beal, Eric Kim, Eric Tzeng +3
Transformers have become the dominant model in natural language processing, owing to their ability to pretrain on massive amounts of data, then transfer to smaller, more specific t…
Bootstrapping Complete The Look at Pinterest
Eileen Li, Eric Kim, Andrew Zhai +2
Putting together an ideal outfit is a process that involves creativity and style intuition. This makes it a particularly difficult task to automate. Existing styling products gener…