Multi-Modality Transformer for E-Commerce: Inferring User Purchase Intention to Bridge the Query-Product Gap
arXiv:2501.14826 · doi:10.1109/BigData62323.2024.10826020
Abstract
E-commerce click-stream data and product catalogs offer critical user behavior insights and product knowledge. This paper propose a multi-modal transformer termed as PINCER, that leverages the above data sources to transform initial user queries into pseudo-product representations. By tapping into these external data sources, our model can infer users' potential purchase intent from their limited queries and capture query relevant product features. We demonstrate our model's superior performance over state-of-the-art alternatives on e-commerce online retrieval in both controlled and real-world experiments. Our ablation studies confirm that the proposed transformer architecture and integrated learning strategies enable the mining of key data sources to infer purchase intent, extract product features, and enhance the transformation pipeline from queries to more accurate pseudo-product representations.
Published in IEEE Big Data Conference 2024, Washington DC
References in corpus (8)
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- Learning Latent Vector Spaces for Product Search
- Pseudo-Relevance Feedback for Multiple Representation Dense Retrieval
- A Transformer-based Embedding Model for Personalized Product Search
- Cross-Market Product Recommendation
- ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
- e-CLIP: Large-Scale Vision-Language Representation Learning in E-commerce
- MAKE: Vision-Language Pre-training based Product Retrieval in Taobao Search