Visual Search at eBay
arXiv:1706.03154 · doi:10.1145/3097983.3098162
Abstract
In this paper, we propose a novel end-to-end approach for scalable visual search infrastructure. We discuss the challenges we faced for a massive volatile inventory like at eBay and present our solution to overcome those. We harness the availability of large image collection of eBay listings and state-of-the-art deep learning techniques to perform visual search at scale. Supervised approach for optimized search limited to top predicted categories and also for compact binary signature are key to scale up without compromising accuracy and precision. Both use a common deep neural network requiring only a single forward inference. The system architecture is presented with in-depth discussions of its basic components and optimizations for a trade-off between search relevance and latency. This solution is currently deployed in a distributed cloud infrastructure and fuels visual search in eBay ShopBot and Close5. We show benchmark on ImageNet dataset on which our approach is faster and more accurate than several unsupervised baselines. We share our learnings with the hope that visual search becomes a first class citizen for all large scale search engines rather than an afterthought.
To appear in 23rd SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2017. A demonstration video can be found at https://youtu.be/iYtjs32vh4g
Cited by in corpus (9)
- Visual Search at Alibaba
- Unsupervised Adversarial Attacks on Deep Feature-based Retrieval with GAN
- Shop The Look: Building a Large Scale Visual Shopping System at Pinterest
- Towards Zero-shot Cross-lingual Image Retrieval
- XCloud: Design and Implementation of AI Cloud Platform with RESTful API Service
- Virtual ID Discovery from E-commerce Media at Alibaba: Exploiting Richness of User Click Behavior for Visual Search Relevance
- Multi-Modal Retrieval using Graph Neural Networks
- Applications of Generative Adversarial Models in Visual Search Reformulation
- Multi-modal dialog for browsing large visual catalogs using exploration-exploitation paradigm in a joint embedding space